Video-ChatGPT is a multimodal artificial intelligence model designed for conversational video understanding and dialogue. It combines a video-adapted visual encoder with a large language model to interpret visual and temporal information across video sequences, enabling it to answer questions, generate summaries, and engage in detailed natural language discussions about dynamic visual content. By integrating spatial and temporal visual features directly into a conversational framework, the system allows users to interactively query, reason about, and analyze complex events, actions, and context depicted within videos.