Video dialogue refers to a multimodal artificial intelligence task and capability where an automated system engages in multi-turn natural language conversation with a user concerning the events, objects, and concepts within a video. Extending beyond single-turn video question answering or static image conversation, video dialogue requires an agent to maintain discourse history while simultaneously processing dynamic spatio-temporal information, visual state transitions, and accompanying audio over time. Systems designed for video dialogue combine computer vision with large language models to temporally ground user queries in relevant video segments, resolve conversational references, and generate coherent, contextually aware responses based on the evolving visual stream.