Hallucination scores are quantitative metrics used to evaluate and measure the likelihood, presence, or severity of fabricated, unfaithful, or factually incorrect content generated by artificial intelligence language models. These values are computed using various techniques, including predictive uncertainty analysis of model tokens, consistency evaluations across repeated generations, and comparison against verified reference data or external knowledge sources. By translating the unreliability or factual inconsistency of generated text into a numerical rating at the token, sentence, or document level, hallucination scores allow practitioners and automated systems to detect false outputs, evaluate model accuracy, and filter or revise untruthful responses.