KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection

Sehyun ChoiTianqing FangZhaowei WangYangqiu Song

article2023EMNLP67 citations

Introduces a plug-and-play decoding framework that combines Monte-Carlo Tree Search with token-level hallucination detection to guide frozen language models toward factually grounded text generation without parameter updates.

Listen

Large language models frequently generate unsupported or non-factual statements, a problem known as hallucination that creates substantial operational, reputational, and compliance risks when deploying artificial intelligence in real-world environments. Standard solutions typically require fine-tuning models on factual reference texts; however, this strategy demands heavy computational expenditures and often causes catastrophic forgetting, where the model loses its general multitasking capabilities.

The article evaluates a plug-and-play decoding framework called Knowledge-Constrained Tree Search (KCTS), designed to guide frozen language models toward generating factual text that strictly aligns with reference knowledge without modifying the underlying model weights.

To achieve this, the researchers developed a tree-search decoding method powered by a novel token-level hallucination detection technique named Reward Inflection Point Approximation (RIPA). RIPA trains a lightweight classifier adding only 0.21% extra parameters to pinpoint the exact token where unsupported claims begin. The search framework uses these scores to simulate potential generation paths and steer the model away from factual errors. The authors validated the system against conventional decoding baselines and large commercial models across two knowledge-intensive tasks: open-domain dialogue using the Wizard of Wikipedia dataset and abstractive news summarization using the CNN/Daily Mail dataset, employing both automated metrics and human evaluations.

The experimental findings show that KCTS substantially increases factual accuracy and grounding across both tasks compared to existing guided decoding methods. On dialogue benchmarks, KCTS achieved a groundedness evaluation score of 91.78 and a knowledge overlap score of 56.06, outperforming previous guided decoding baselines. In summarization tasks, it delivered significant improvements in factual consistency metrics and token overlap. Human evaluators consistently rated responses generated by KCTS higher in groundedness, fluency, and relevance than standard decoding baselines, while confirming that the system generates original phrasing rather than simply copying reference text verbatim. Furthermore, the analysis demonstrated that applying the search constraint to only the initial 16 to 32 tokens successfully anchors the factual trajectory of the entire response.

These results indicate that enterprises can mitigate hallucination risks and enforce factual compliance without incurring the high financial and computational costs of continuous model retraining. Because the approach keeps the core language model frozen, organizations can preserve broad multi-task versatility while applying reliable factual controls as a modular add-on.

For practical implementation, organizations should consider adopting knowledge-constrained decoding layers to safeguard critical generation workflows. Engineering teams should leverage the option to constrain only initial prefix tokens or adjust simulation counts, allowing them to balance groundedness against computational speed. Future development should focus on integrating automated knowledge retrieval pipelines to evaluate the framework under realistic end-to-end information retrieval settings.

The primary limitation of this method is increased latency and computational cost per generated token due to the multiple simulation passes required during tree search. Additionally, the system enforces faithfulness to the provided input knowledge rather than external ground truth, meaning incorrect reference data will result in faithfully reproduced inaccuracies. Confidence in the demonstrated factual improvements is high, though careful latency engineering is necessary before deploying in time-sensitive production environments.

Choi et al (2023).pdf
  • Paper: Controlled Decoding from Language Models, Sidharth Mudgal et al. (2024). Controlled Decoding extends inference-time steering of frozen language models with prefix-level reward scoring, offering a natural next step from KCTS’s knowledge-guided search.
Cover for KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection

Abstract

Large Language Models (LLMs) have demonstrated remarkable human-level natural language generation capabilities. However, their potential to generate misinformation, often called the hallucination problem, poses a significant risk to their deployment. A common approach to address this issue is to retrieve relevant knowledge and fine-tune the LLM with the knowledge in its input. Unfortunately, this method incurs high training costs and may cause catastrophic forgetting for multi-tasking models. To overcome these limitations, we propose a knowledge-constrained decoding method called KCTS (Knowledge-Constrained Tree Search), which guides a frozen LM to generate text aligned with the reference knowledge at each decoding step using a knowledge classifier score and MCTS (Monte-Carlo Tree Search). To adapt the sequence-level knowledge classifier to token-level guidance, we also propose a novel token-level hallucination detection method called RIPA (Reward Inflection Point Approximation). Our empirical results on knowledge-grounded dialogue and abstractive summarization demonstrate the strength of KCTS^1 as a plug-and-play, model-agnostic decoding method that can effectively reduce hallucinations in natural language generation.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Work
  • 3 Problem Statement
  • 4 The KCTS Method
  • 4.1 Monte-Carlo Tree Search Decoding
  • 4.2 Token-Level Hallucination Detection
  • 4.3 Knowledge Weighted Decoding (KWD)
  • 5 Experiments Setup
  • 5.1 Datasets
  • 5.2 Evaluation Metrics
  • 5.3 Baselines
  • 5.4 Implementation Details
  • 6 Main Evaluation
  • 6.1 KGD Results
  • 6.2 Summarization Results
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Instruction Templates
  • B Application to LLMs
  • C Implementation Detail
  • D Analysis on Partial Hallucination Data
  • E Classifier Performance
  • F Generated Examples

Knowls

  1. Knowl 1 — Knowledge-constrained generation objective

    model/method

    Let xx denote the task prompt and any input context, kk the reference knowledge that a response must follow, and y=(y1,…,yT)y=(y_1,\ldots,y_T) the generated token sequence. Let αk∈{0,1}\alpha_k\in\{0,1\} indicate whether a response is grounded in kk, and define its groundedness probability as f(y,k)=P(αk=1∣y,k)f(y,k)=P(\alpha_k=1\mid y,k). The paper formulates constrained generation as favoring outputs that are both probable under the language model and grounded in the reference knowledge:

    PLM(y∣x,k,αk)∝PLM(y∣x)f(y,k),y∗=arg⁡max⁡yPLM(y∣x)f(y,k).P_{\mathrm{LM}}(y\mid x,k,\alpha_k)\propto P_{\mathrm{LM}}(y\mid x)f(y,k), \qquad y^*=\arg\max_y P_{\mathrm{LM}}(y\mid x)f(y,k).

    For autoregressive decoding, the token choice at position tt is guided by the language model probability of the next token and the groundedness score for the resulting partial sequence, f(y≤t,k)f(y_{\le t},k). Because groundedness is most directly defined for completed responses, a prefix score f(y<t,k)f(y_{<t},k) is treated as an approximation to the groundedness of a future continuation sampled from the language model. Groundedness here means support by the supplied kk; it does not by itself establish that kk is factually correct.

  2. Knowl 2 — KCTS uses Monte Carlo tree search to guide token selection

    algorithm

    Knowledge-Constrained Tree Search (KCTS) uses a groundedness classifier to guide a frozen autoregressive language model without changing the base model's generation objective through fine-tuning. Its input at each decoding step is the current generated prefix, the task context xx, the reference knowledge kk, the language model, and a classifier that estimates groundedness for partial sequences; its output is the next committed token.

    KCTS treats the current prefix as the root of a search tree and each child as a candidate next token. In each simulation it selects a path from the root using PUCT: the selection score balances the node's estimated groundedness reward against an exploration term that uses the language model's probability for that token and the parent and child visit counts. If the selected leaf is not an end-of-sequence token, KCTS expands it with the top-kk next-token candidates from the language model. Instead of completing a rollout to the end of the response, it evaluates the selected prefix directly with the token-level groundedness classifier. It propagates that score back along the path, updating the path's value estimates by mean aggregation over simulations.

    After the preset number of simulations, KCTS commits the root child with the highest visit count as the next output token, makes that child the new root, and repeats until an end-of-sequence token or the generation limit is reached. The reported experiments use 5050 simulations per token, PUCT exploration constant cpuct=3c_{\mathrm{puct}}=3, and top-5050 token filtering; the MCTS decoding also uses a repetition penalty of 1.21.2. The method's central distinction from greedy weighted decoding is that token commitment depends on simulated future search and visit counts, not only on a locally reweighted token score.

  3. Knowl 3 — RIPA assigns token labels at the onset of hallucination

    model/method

    Reward Inflection Point Approximation (RIPA) adapts a sequence-level groundedness classifier to provide guidance for partial generations. Given a response with token-level annotations, RIPA labels tokens before the first unsupported or hallucinated token as grounded (label 11), and labels the first hallucinated token and all subsequent tokens as ungrounded (label 00). Thus, the token labels change at the hallucination onset rather than inheriting a single label from the entire response.

    The classifier trained on these labels estimates whether a prefix is expected to lead to a grounded continuation. The design is intended to avoid labeling benign material before a later hallucination as ungrounded, a problem the paper identifies for methods that assign one sequence label to every token. RIPA also labels later prefixes after hallucination as ungrounded, so search methods such as KCTS can reduce exploration beneath those prefixes. The paper hypothesizes that this hallucination-onset signal is a more useful approximation of future groundedness; it does not claim that every individual onset is detected perfectly.

  4. Knowl 4 — Synthetic data supplies RIPA's token-level supervision

    model/method

    Because fine-grained human labels for hallucination onset are difficult to obtain, the paper creates synthetic negative examples using knowledge shuffling and partial hallucination. For knowledge shuffling, it takes a training example consisting of response yy, context xx, and knowledge kk, substitutes a randomly selected other knowledge item k′k', and labels all response tokens ungrounded. The response and context remain unchanged, while the substituted knowledge is expected to make the response unsupported by the supplied evidence.

    For partial hallucination, the procedure also substitutes a knowledge item, samples an interior token position ii, retains the original response prefix through ii, and asks a language model to complete it conditioned on the context and mismatched knowledge. Only the generated completion is labeled ungrounded, leaving the preceding response tokens as the grounded portion. The completion is sampled at temperature greater than 11 to encourage unsupported continuations; the experiments use temperature 1.41.4.

    For Wizard of Wikipedia dialogue, the training data contained 20,000 original positive examples, 10,000 knowledge-shuffle negatives, and 8,832 partial-hallucination negatives, giving 20,000 positive and 18,832 negative examples. For CNN/DailyMail summarization, where the document itself serves as the knowledge and context distinction is not applicable, the authors used partial hallucination only: 13,180 positive examples and 12,811 synthetic negatives. The generated datasets were randomly split 9:1 into training and test portions.

  5. Knowl 5 — Classifier implementation and Knowledge Weighted Decoding

    model/method

    The groundedness classifier is trained as a binary sequence classifier using the language model's decoder representations and a linear classification layer. The base language-model weights remain unchanged; lightweight LoRA adapters are applied to decoder layers for classifier training, and the added trainable weights amount to 0.21%0.21\% of the model weights. The adapters can be enabled or disabled at inference, allowing the same base model to serve as the generator.

    The paper also introduces Knowledge Weighted Decoding (KWD), which uses the RIPA classifier in a FUDGE-style weighted decoding procedure rather than tree search. KWD obtains candidate tokens from the language model, scores candidate continuations with RIPA, and re-ranks or reweights the candidates using those token-level groundedness estimates. KWD therefore tests the value of RIPA guidance independently of KCTS's MCTS search.

  6. Knowl 6 — Evaluation setting for knowledge-constrained decoding

    experimental setup

    The experiments evaluate knowledge-constrained decoding on two tasks with reference knowledge supplied to the generator. Knowledge-grounded dialogue uses the unseen-topic test portion of Wizard of Wikipedia, with dialogue history as context and the dataset's gold knowledge as the evidence. Abstractive summarization uses the CNN/DailyMail test set, treating the source article as the knowledge that the summary should follow. The main guided-decoding experiments use Flan-T5-XL as the base model.

    Comparisons include FUDGE, NADO, and MCTS decoding baselines, as well as KWD and KCTS. The paper also reports ChatGPT, GPT-3.5, and other instruction-tuned models for reference, but cautions that they are not directly comparable because of differences in model size and training. Automatic evaluation covers knowledge overlap (Knowledge-F1 and K-Copy), reference-token overlap (including BLEU, ROUGE-L, ChrF, and METEOR), and learned evaluators. Dialogue evaluation includes naturalness, coherence, and groundedness; summarization evaluation includes coherence, consistency, fluency, relevance, and MFMA. For decoding, the response limits are 32 tokens for dialogue and 64 for summarization, with top-p=0.95p=0.95, temperature 11, and top-5050 filtering.

  7. Knowl 7 — Dialogue results show stronger groundedness with RIPA-guided decoding

    data/table

    On the Wizard of Wikipedia unseen-topic test set, the directly comparable decoding methods use Flan-T5-XL. In the metric order Knowledge-F1 (KF1), K-Copy, token-overlap F1, BLEU, ROUGE-L, ChrF, METEOR, UniEval naturalness (N), coherence (C), groundedness (G), and the reported classifier score ff, their results are:

    • Zero-shot Flan-T5-XL: 34.50,37.07,21.18,6.81,19.64,24.88,18.53,71.69,82.21,75.70,88.7534.50, 37.07, 21.18, 6.81, 19.64, 24.88, 18.53, 71.69, 82.21, 75.70, 88.75.
    • FUDGE: 55.30,54.04,29.43,11.72,27.35,31.50,26.00,73.68,88.20,83.53,94.5455.30, 54.04, 29.43, 11.72, 27.35, 31.50, 26.00, 73.68, 88.20, 83.53, 94.54.
    • NADO: 50.20,50.10,27.86,10.57,26.01,29.84,24.51,74.14,88.35,81.10,92.7650.20, 50.10, 27.86, 10.57, 26.01, 29.84, 24.51, 74.14, 88.35, 81.10, 92.76.
    • MCTS: 55.54,54.21,29.56,11.69,27.48,31.60,26.08,74.54,88.16,83.90,95.0755.54, 54.21, 29.56, 11.69, 27.48, 31.60, 26.08, 74.54, 88.16, 83.90, 95.07.
    • KWD: 58.19,56.58,30.71,12.74,28.27,33.40,28.10,70.27,90.51,87.86,97.5458.19, 56.58, 30.71, 12.74, 28.27, 33.40, 28.10, 70.27, 90.51, 87.86, 97.54.
    • KCTS: 56.06,51.90,30.54,11.42,27.43,35.22,28.92,62.32,92.78,91.78,98.3056.06, 51.90, 30.54, 11.42, 27.43, 35.22, 28.92, 62.32, 92.78, 91.78, 98.30.

    Higher is preferred for all listed metrics except K-Copy, where lower is preferred. Among the compared guided methods, KWD has the highest KF1 and token-overlap F1, while KCTS has the best K-Copy, ChrF, METEOR, UniEval coherence and groundedness, and reported ff score. KCTS's naturalness score is lower than those of the decoding baselines, so the groundedness improvements do not imply improvement on every dimension.

  8. Knowl 8 — Summarization results favor KCTS on overlap and MFMA

    data/table

    On CNN/DailyMail, guided decoding uses Flan-T5-XL as the base model. In the metric order KF1, K-Copy, token-overlap F1, BLEU, ROUGE-L, ChrF, METEOR, UniEval coherence, consistency, fluency, relevance, and MFMA score, the reported results are:

    • Flan-T5-XL: 17.04,10.18,32.21,8.74,24.02,30.27,24.47,84.82,86.02,89.90,81.28,64.5517.04, 10.18, 32.21, 8.74, 24.02, 30.27, 24.47, 84.82, 86.02, 89.90, 81.28, 64.55.
    • FUDGE: 18.68,10.70,33.51,9.32,24.83,31.06,24.93,90.52,90.61,83.37,82.00,71.3518.68, 10.70, 33.51, 9.32, 24.83, 31.06, 24.93, 90.52, 90.61, 83.37, 82.00, 71.35.
    • NADO: 20.35,11.72,35.10,10.93,26.22,33.50,27.34,92.26,93.72,88.41,84.49,72.0120.35, 11.72, 35.10, 10.93, 26.22, 33.50, 27.34, 92.26, 93.72, 88.41, 84.49, 72.01.
    • MCTS: 17.86,10.04,34.59,9.00,25.85,30.90,25.12,94.30,94.28,86.51,85.90,71.2817.86, 10.04, 34.59, 9.00, 25.85, 30.90, 25.12, 94.30, 94.28, 86.51, 85.90, 71.28.
    • KWD: 20.39,11.63,36.24,12.30,27.20,34.25,28.46,96.24,96.64,91.60,88.48,85.1120.39, 11.63, 36.24, 12.30, 27.20, 34.25, 28.46, 96.24, 96.64, 91.60, 88.48, 85.11.
    • KCTS: 22.97,13.29,38.27,14.21,28.10,37.18,31.37,95.85,96.03,90.24,87.16,85.3622.97, 13.29, 38.27, 14.21, 28.10, 37.18, 31.37, 95.85, 96.03, 90.24, 87.16, 85.36.

    Higher is preferred except for K-Copy. KCTS has the highest KF1, BLEU, ROUGE-L, ChrF, METEOR, and MFMA score among these guided methods. KWD is higher than KCTS on all four reported UniEval dimensions, while KCTS has higher knowledge overlap; the authors note that this difference may affect the learned evaluator's assessment of abstractive summaries.

  9. Knowl 9 — Grounded initial tokens can improve later dialogue generation

    data/table

    The paper tests whether KCTS estimates future groundedness by constraining only the first TT generated tokens with KCTS and then letting the base language model complete the response using nucleus sampling. For T=5,10,16,32T=5,10,16,32, the reported metric order is KF1, K-Copy, BLEU, ROUGE-L, UniEval coherence (C), groundedness (G), and ff:

    • T=5T=5: 48.78,48.22,10.17,25.39,90.58,85.87,90.5848.78, 48.22, 10.17, 25.39, 90.58, 85.87, 90.58.
    • T=10T=10: 48.24,48.05,9.98,25.87,90.22,86.41,85.4348.24, 48.05, 9.98, 25.87, 90.22, 86.41, 85.43.
    • T=16T=16: 51.49,48.67,11.07,26.44,92.83,89.99,92.7651.49, 48.67, 11.07, 26.44, 92.83, 89.99, 92.76.
    • T=32T=32: 56.06,51.90,11.42,27.43,92.78,91.78,98.3056.06, 51.90, 11.42, 27.43, 92.78, 91.78, 98.30.

    The general increase in groundedness-related scores as more initial tokens are constrained supports the authors' claim that the selected prefix can steer the unconstrained continuation toward more grounded responses. The non-monotonic scores at T=10T=10 show that this is not a strictly increasing trend for every metric. Choosing TT also offers a performance-versus-decoding-cost trade-off.

  10. Knowl 10 — Human evaluation favors KCTS over decoding baselines

    data/table

    Three evaluators used a 3-point Likert scale to assess generated text. For Wizard of Wikipedia, 100 examples were rated for fluency, relevance, and groundedness; the reported mean scores in that order were ChatGPT 3.00,2.73,2.623.00, 2.73, 2.62, Flan-T5-XL 2.64,2.30,1.952.64, 2.30, 1.95, FUDGE 2.82,2.35,2.192.82, 2.35, 2.19, and KCTS 2.92,2.55,2.372.92, 2.55, 2.37. The corresponding K-Copy rates were 0.04,0.12,0.21,0.04, 0.12, 0.21, and 0.170.17. On examples judged not to copy the knowledge, KCTS scored 2.91,2.61,2.242.91, 2.61, 2.24 for fluency, relevance, and groundedness, compared with FUDGE's 2.78,2.40,1.972.78, 2.40, 1.97.

    For CNN/DailyMail, 50 examples were rated for fluency, groundedness, and completeness. Mean scores were ChatGPT 3.00,2.93,2.883.00, 2.93, 2.88, Flan-T5-XL 2.81,2.60,2.132.81, 2.60, 2.13, FUDGE 2.89,2.90,2.312.89, 2.90, 2.31, and KCTS 2.95,2.97,2.402.95, 2.97, 2.40. Thus KCTS exceeded the decoding baselines on all three summarization dimensions and on dialogue groundedness and relevance, while ChatGPT retained the highest dialogue fluency and relevance. The reported inter-rater Krippendorff alpha values were 0.57,0.46,0.77,0.310.57, 0.46, 0.77, 0.31 for the four dialogue judgments and 0.35,0.44,0.190.35, 0.44, 0.19 for summarization fluency, groundedness, and completeness, respectively.

  11. Knowl 11 — RIPA performs best among evaluated token-level classifiers

    empirical result

    On held-out portions of the synthetic classifier datasets for Wizard of Wikipedia and CNN/DailyMail, the sequence-level groundedness classifier achieved very high accuracy. RIPA's accuracy closely followed the sequence-level classifier and was higher than the evaluated token-level alternatives: random input truncation, used by FUDGE, and uniform sequence-label assignment to tokens, used by NADO. The paper attributes the weaker performance of random truncation partly to noisy negative labels when a truncated prefix no longer contains the hallucination. It attributes the token-labeling baseline's weaker performance partly to labeling benign pre-hallucination tokens as negative. The article reports these comparisons graphically rather than giving exact classifier-accuracy values.

    The paper also checks the synthetic partial-hallucination examples with UniEval: groundedness for dialogue falls from 0.950.95 on original examples to 0.630.63 on generated partial-hallucination examples, and summarization consistency falls from 0.880.88 to 0.360.36. Of the tokens sampled in the synthetic completions, 78%78\% for dialogue and 68%68\% for summarization are among the base model's top-5050 tokens. These checks support, but do not prove, that the synthetic negatives reflect both reduced groundedness and token choices relevant to the decoding search space.

  12. Knowl 12 — Knowledge-constrained decoding trades inference cost for grounding

    limitation

    KCTS increases decoding computation because it performs NN language-model and discriminator forward passes for each generated token to run NN MCTS simulations; the experiments use N=50N=50. The paper contrasts this with NADO, which requires two forward passes per token, and FUDGE, which requires discriminator evaluations for its top-kk candidates. Reducing the number of simulations or candidate tokens can reduce cost, but the paper frames generation speed and groundedness as a trade-off. It suggests smaller discriminator weights or early stopping of MCTS as possible efficiency improvements.

    The experiments use gold knowledge and do not study retrieval errors, so performance when knowledge must be retrieved remains untested. KCTS constrains outputs to the supplied reference; it does not independently verify that the reference itself is true, and an incorrect knowledge source can therefore still lead to misinformation. The paper also cautions that comparisons with ChatGPT are not necessarily fair because its training data, architecture, and mechanisms are undisclosed.

Coverage note — Illustrative generated examples and the preliminary GPT-3.5 Pre-KWD experiment are omitted because they are secondary case studies rather than load-bearing contributions to the main method or its primary evaluation.

References

  1. 1.Renat Aksitov, Chung-Ching Chang, David Reitter, Siamak Shakeri, and Yunhsuan Sung. 2023. Characterizing attribution and fluency tradeoffs for retrieval-augmented large language models.
  2. 2.Kushal Arora, Kurt Shuster, Sainbayar Sukhbaatar, and Jason Weston. 2022. Director: Generator-classifiers for supervised language modeling. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 512–526, Online only. Association for Computational Linguistics.
  3. 3.Akari Asai, Sewon Min, Zexuan Zhong, and Danqi Chen. 2023. Retrieval-based language models and applications. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 6: Tutorial Abstracts), pages 41–46, Toronto, Canada. Association for Computational Linguistics.
  4. 4.Hendrik Baier and Mark H. M. Winands. 2012. Time management for monte-carlo tree search in go. In Advances in Computer Games, pages 39–51, Berlin, Heidelberg. Springer Berlin Heidelberg.
  5. 5.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan. Association for Computational Linguistics.
  6. 6.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity.
  7. 7.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego De Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack Rae, Erich Elsen, and Laurent Sifre. 2022. Improving language models by retrieving from trillions of tokens. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 2206–2240. PMLR.
  8. 8.Antoine Chaffin, Vincent Claveau, and Ewa Kijak. 2022a. PPL-MCTS: Constrained textual generation through discriminator-guided MCTS decoding. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2953–2967, Seattle, United States. Association for Computational Linguistics.
  9. 9.Antoine Chaffin, Thomas Scialom, Sylvain Lamprier, Jacopo Staiano, Benjamin Piwowarski, Ewa Kijak, and Vincent Claveau. 2022b. Which discriminator for cooperative text generation? In SIGIR ’22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, July 11 - 15, 2022, pages 2360–2365. ACM.
  10. 10.Chung-Ching Chang, David Reitter, Renat Aksitov, and Yun-Hsuan Sung. 2023. Kl-divergence guided temperature sampling.
  11. 11.Shiqi Chen, Yiran Zhao, Jinghan Zhang, I-Chun Chern, Siyang Gao, Pengfei Liu, and Junxian He. 2023. Felm: Benchmarking factuality evaluation of large language models.
  12. 12.Jonathan Chevelu, Thomas Lavergne, Yves Lepage, and Thierry Moudenc. 2009. Introduction of a new paraphrase generation tool based on Monte-Carlo sampling. In Proceedings of the ACL-IJCNLP 2009 Conference Short Papers, pages 249–252, Suntec, Singapore. Association for Computational Linguistics.
  13. 13.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models.
  14. 14.Rémi Coulom. 2007. Efficient selectivity and backup operators in monte-carlo tree search. In Computers and Games, pages 72–83, Berlin, Heidelberg. Springer Berlin Heidelberg.
  15. 15.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations.
  16. 16.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022a. Llm.int8(): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339.
  17. 17.Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 2022b. 8-bit optimizers via block-wise quantization. 9th International Conference on Learning Representations, ICLR.
  18. 18.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019. Wizard of Wikipedia: Knowledge-powered conversational agents. In Proceedings of the International Conference on Learning Representations (ICLR).
  19. 19.Xidong Feng, Ziyu Wan, Muning Wen, Ying Wen, Weinan Zhang, and Jun Wang. 2023. Alphazero-like tree-search can guide large language model decoding and training. arXiv preprint arXiv:2309.17179.
  20. 20.Thibault Févry, Livio Baldini Soares, Nicholas Fitzgerald, Eunsol Choi, and Tom Kwiatkowski. 2020. Entities as experts: Sparse memory access with entity supervision. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4937–4951, Online. Association for Computational Linguistics.
  21. 21.Robert M. French. 1999. Catastrophic forgetting in connectionist networks. Trends in Cognitive Sciences, 3(4):128–135.
  22. 22.Saibo Geng, Martin Josifosky, Maxime Peyrard, and Robert West. 2023. Flexible grammar-based constrained decoding for language models.
  23. 23.Hangfeng He, Hongming Zhang, and Dan Roth. 2022. Rethinking with retrieval: Faithful large language model inference.
  24. 24.Tianxing He, Jun Liu, Kyunghyun Cho, Myle Ott, Bing Liu, James Glass, and Fuchun Peng. 2021. Analyzing the forgetting problem in pretrain-finetuning of open-domain dialogue response models. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1121–1133, Online. Association for Computational Linguistics.
  25. 25.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models.
  26. 26.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880, Online. Association for Computational Linguistics.
  27. 27.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023a. Survey of hallucination in natural language generation. ACM Comput. Surv., 55(12).
  28. 28.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023b. Survey of hallucination in natural language generation. ACM Comput. Surv., 55(12).
  29. 29.Ziwei Ji, Zihan Liu, Nayeon Lee, Tiezheng Yu, Bryan Wilie, Min Zeng, and Pascale Fung. 2023c. RHO: Reducing hallucination in open-domain dialogues with knowledge grounding. In Findings of the Association for Computational Linguistics: ACL 2023, pages 4504–4522, Toronto, Canada. Association for Computational Linguistics.
  30. 30.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  31. 31.Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation.
  32. 32.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 114(13):3521–3526.
  33. 33.Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2021. GeDi: Generative discriminator guided sequence generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4929–4952, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  34. 34.Klaus Krippendorff. 2011. Computing krippendorff’s alpha-reliability.
  35. 35.Sachin Kumar, Biswajit Paria, and Yulia Tsvetkov. 2022. Gradient-based constrained sampling from language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2251–2277, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  36. 36.Sylvain Lamprier, Thomas Scialom, Antoine Chaffin, Vincent Claveau, Ewa Kijak, Jacopo Staiano, and Benjamin Piwowarski. 2022. Generative cooperative networks for natural language generation. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 11891–11905. PMLR.
  37. 37.Rémi Leblond, Jean-Baptiste Alayrac, Laurent Sifre, Miruna Pislar, Lespiau Jean-Baptiste, Ioannis Antonoglou, Karen Simonyan, and Oriol Vinyals. 2021. Machine translation decoding beyond beam search. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8410–8434, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  38. 38.Hwanhee Lee, Kang Min Yoo, Joonsuk Park, Hwaran Lee, and Kyomin Jung. 2022. Masked summarization to generate factually inconsistent summaries for improved factual consistency checking. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 1019–1030, Seattle, United States. Association for Computational Linguistics.
  39. 39.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems, volume 33, pages 9459–9474. Curran Associates, Inc.
  40. 40.Rongzhong Lian, Min Xie, Fan Wang, Jinhua Peng, and Hua Wu. 2019. Learning to select knowledge for response generation in dialog systems. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19, pages 5081–5087. International Joint Conferences on Artificial Intelligence Organization.
  41. 41.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  42. 42.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics.
  43. 43.Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021. DExperts: Decoding-time controlled text generation with experts and anti-experts. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6691–6706, Online. Association for Computational Linguistics.
  44. 44.Jiacheng Liu, Andrew Cohen, Ramakanth Pasunuru, Yejin Choi, Hannaneh Hajishirzi, and Asli Celikyilmaz. 2023a. Making ppo even better: Value-guided monte-carlo tree search decoding.
  45. 45.Nelson F. Liu, Tianyi Zhang, and Percy Liang. 2023b. Evaluating verifiability in generative search engines.
  46. 46.Xin Liu, Muhammad Khalifa, and Lu Wang. 2023c. BOLT: Fast energy-based controlled text generation with tunable biases. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 186–200, Toronto, Canada. Association for Computational Linguistics.
  47. 47.Ximing Lu, Sean Welleck, Peter West, Liwei Jiang, Jungo Kasai, Daniel Khashabi, Ronan Le Bras, Lianhui Qin, Youngjae Yu, Rowan Zellers, Noah A. Smith, and Yejin Choi. 2022. NeuroLogic a*esque decoding: Constrained text generation with lookahead heuristics. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 780–799, Seattle, United States. Association for Computational Linguistics.
  48. 48.Ximing Lu, Peter West, Rowan Zellers, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. NeuroLogic decoding: (un)supervised neural text generation with predicate logic constraints. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4288–4299, Online. Association for Computational Linguistics.
  49. 49.Potsawee Manakul, Adian Liusie, and Mark J. F. Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models.
  50. 50.Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, and Sayak Paul. 2022. Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft.
  51. 51.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  52. 52.Tao Meng, Sidi Lu, Nanyun Peng, and Kai-Wei Chang. 2022. Controllable text generation with neurally-decomposed oracle. In Advances in Neural Information Processing Systems, volume 35, pages 28125–28139. Curran Associates, Inc.
  53. 53.Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, Edouard Grave, Yann LeCun, and Thomas Scialom. 2023. Augmented language models: a survey.
  54. 54.Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. Factscore: Fine-grained atomic evaluation of factual precision in long form text generation.
  55. 55.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback.
  56. 56.Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4812–4829, Online. Association for Computational Linguistics.
  57. 57.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  58. 58.Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, and Jianfeng Gao. 2023. Check your facts and try again: Improving large language models with external knowledge and automated feedback.
  59. 59.Maja Popovic. 2015. chrF: character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395, Lisbon, Portugal. Association for Computational Linguistics.
  60. 60.Lianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula, Jena D. Hwang, Ronan Le Bras, Antoine Bosselut, and Yejin Choi. 2020. Back to the future: Unsupervised backprop-based decoding for counterfactual and abductive commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 794–805, Online. Association for Computational Linguistics.
  61. 61.Lianhui Qin, Sean Welleck, Daniel Khashabi, and Yejin Choi. 2022. COLD decoding: Energy-based constrained text generation with langevin dynamics. In Advances in Neural Information Processing Systems.
  62. 62.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  63. 63.Hannah Rashkin, David Reitter, Gaurav Singh Tomar, and Dipanjan Das. 2021. Increasing faithfulness in knowledge-grounded dialogue with controllable features. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 704–718, Online. Association for Computational Linguistics.
  64. 64.Vikas Raunak, Arul Menezes, and Marcin Junczys-Dowmunt. 2021. The curious case of hallucinations in neural machine translation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1172–1183, Online. Association for Computational Linguistics.
  65. 65.Stephen Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends in Information Retrieval, 3:333–389.
  66. 66.Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018. Object hallucination in image captioning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4035–4045, Brussels, Belgium. Association for Computational Linguistics.
  67. 67.Christopher D. Rosin. 2011. Multi-armed bandits with episode context. Annals of Mathematics and Artificial Intelligence, 61(3):203–230.
  68. 68.Ohad Rubin and Jonathan Berant. 2023. Long-range language modeling with self-retrieval.
  69. 69.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. 2022. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations.
  70. 70.Thomas Scialom, Paul-Alexis Dray, Jacopo Staiano, Sylvain Lamprier, and Benjamin Piwowarski. 2021. To beam or not to beam: That is a question of cooperation for language gans. In Advances in Neural Information Processing Systems, volume 34, pages 26585–26597. Curran Associates, Inc.
  71. 71.Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get to the point: Summarization with pointer-generator networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1073–1083, Vancouver, Canada. Association for Computational Linguistics.
  72. 72.Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen tau Yih. 2023. Replug: Retrieval-augmented black-box language models.
  73. 73.Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3784–3803, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  74. 74.David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. 2017. Mastering the game of go without human knowledge. Nature, 550(7676):354–359.
  75. 75.Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2022. One embedder, any task: Instruction-finetuned text embeddings.
  76. 76.Yi Tay. 2023. A new open source flan 20b with ul2.
  77. 77.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819, New Orleans, Louisiana. Association for Computational Linguistics.
  78. 78.Pat Verga, Haitian Sun, Livio Baldini Soares, and William Cohen. 2021. Adaptable and interpretable neural MemoryOver symbolic knowledge. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3678–3691, Online. Association for Computational Linguistics.
  79. 79.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020, Online. Association for Computational Linguistics.
  80. 80.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit, and Xudong Shen. 2022a. Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5085–5109, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  81. 81.Zhaowei Wang, Hongming Zhang, Tianqing Fang, Yangqiu Song, Ginny Wong, and Simon See. 2022b. Subeventwriter: Iterative sub-event sequence generation with coherence controller. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1590–1604.
  82. 82.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  83. 83.Kaiqiang Xu, Xinchen Wan, Hao Wang, Zhenghang Ren, Xudong Liao, Decang Sun, Chaoliang Zeng, and Kai Chen. 2021. Tacc: A full-stack cloud computing infrastructure for machine learning tasks. arXiv preprint arXiv:2110.01556.
  84. 84.Yan Xu, Deqian Kong, Dehong Xu, Ziwei Ji, Bo Pang, Pascale Fung, and Ying Nian Wu. 2023. Diverse and faithful knowledge-grounded dialogue generation via sequential posterior inference. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org.
  85. 85.Kevin Yang and Dan Klein. 2021. FUDGE: Controlled text generation with future discriminators. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3511–3535, Online. Association for Computational Linguistics.
  86. 86.Hugh Zhang, Daniel Duckworth, Daphne Ippolito, and Arvind Neelakantan. 2021. Trading off diversity and quality in natural language generation. In Proceedings of the Workshop on Human Evaluation of NLP Systems (HumEval), pages 25–33, Online. Association for Computational Linguistics.
  87. 87.Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A. Smith. 2023. How language model hallucinations can snowball.
  88. 88.Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022a. Towards a unified multi-dimensional evaluator for text generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2023–2038, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  89. 89.Zexuan Zhong, Tao Lei, and Danqi Chen. 2022b. Training language models with memory augmentation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5657–5673, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.

Citation

MLA
Choi, S., et al. “KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 14035–53, https://doi.org/10.18653/v1/2023.emnlp-main.867.
APA
Choi, S., Fang, T., Wang, Z., & Song, Y. (2023). KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 14035–14053. https://doi.org/10.18653/v1/2023.emnlp-main.867
Chicago
Choi, S., T. Fang, Z. Wang, and Y. Song. 2023. “KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 14035–53. https://doi.org/10.18653/v1/2023.emnlp-main.867.
Harvard
Choi, S. et al. (2023) “KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 14035–14053. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.867.
Vancouver
1. Choi S, Fang T, Wang Z, Song Y (2023) KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 14035–14053

BibTeX

@inproceedings{choi-etal-2023-kcts,
    title = "{KCTS}: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection",
    author = "Choi, Sehyun  and
      Fang, Tianqing  and
      Wang, Zhaowei  and
      Song, Yangqiu",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.867/",
    doi = "10.18653/v1/2023.emnlp-main.867",
    pages = "14035--14053"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/