Built independently by an author, for readers. Read the story and support ChapterPal

keyword

zero-resource black-box hallucination detection

Zero-resource black-box hallucination detection is a method for identifying factual errors and fabricated statements in text generated by large language models without using external reference databases and without requiring access to internal model parameters or output probability distributions. In this framework, the black-box designation indicates that the evaluation process operates entirely through standard text inputs and outputs, making it suitable for proprietary or closed-source models where internal states and token probabilities are inaccessible. The zero-resource component signifies that the system does not depend on external knowledge bases, search engines, or curated fact-checking datasets. Instead, this approach evaluates factuality by measuring the internal consistency of multiple stochastic responses generated from the same prompt, leveraging the tendency of models to generate consistent facts when grounded in learned knowledge and contradictory or divergent details when fabricating information.

1 item

SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

Potsawee Manakul, Adian Liusie, Mark J. F. Gales

OrganizationsALTA InstituteUniversity of Cambridge

Why you should read this

Proposes SelfCheckGPT, a zero-resource sampling method that detects hallucinations in black-box large language models by measuring consistency across stochastically generated responses without requiring external databases or token probability distributions.

Generative Large Language Models (LLMs) such as GPT-3 are capable of generating highly fluent responses to a wide variety of user prompts. However, LLMs are known to hallucinate facts and make non-factual statements which can undermine trust in their output. Existing fact-checking approaches either require access to the output probability distribution (which may not be available for systems such as ChatGPT) or external databases that are interfaced via separate, often complex, modules. In this work, we propose "SelfCheckGPT", a simple sampling-based approach that can be used to fact-check the responses of black-box models in a zero-resource fashion, i.e. without an external database. SelfCheckGPT leverages the simple idea that if an LLM has knowledge of a given concept, sampled responses are likely to be similar and contain consistent facts. However, for hallucinated facts, stochastically sampled responses are likely to diverge and contradict one another. We investigate this approach by using GPT-3 to generate passages about individuals from the WikiBio dataset, and manually annotate the factuality of the generated passages. We demonstrate that SelfCheckGPT can: i) detect non-factual and factual sentences; and ii) rank passages in terms of factuality. We compare our approach to several baselines and show that our approach has considerably higher AUC-PR scores in sentence-level hallucination detection and higher correlation scores in passage-level factuality assessment compared to grey-box methods.

Added

2026-09-28