Image-text alignment scores are quantitative metrics used to evaluate how accurately and faithfully the visual content of an image corresponds to the semantic meaning of an accompanying text description or prompt. In multimodal artificial intelligence and text-to-image synthesis, these scores assess whether key elements specified in the text, such as objects, attributes, actions, and spatial relationships, are correctly depicted in the visual output. They are commonly determined through human judgment protocols or automated scoring models, such as vision-language networks that calculate the cosine similarity between text and image embeddings in a shared representation space. These scores provide a standardized way to measure semantic fidelity independently of general visual quality metrics, helping researchers evaluate and compare how well generative models follow natural language instructions.