Visual question answering models are multimodal artificial intelligence systems designed to process visual inputs, such as images or videos, and generate accurate natural language answers to questions about that visual content. These models bridge computer vision and natural language processing by extracting visual features, parsing question semantics, and performing joint reasoning over both modalities. Depending on their architecture, VQA models can identify objects, analyze spatial relationships, recognize and interpret embedded text, and leverage contextual knowledge to answer a wide variety of queries.