Multimodal dialogues are conversational interactions between participants, such as humans and artificial intelligence systems, that incorporate and exchange information across multiple modes of communication rather than relying solely on text. In these multi-turn exchanges, participants can combine, alternate, and interpret various modalities, including spoken language, written text, visual imagery, video, and audio, as both inputs and outputs. Systems supporting multimodal dialogue dynamically track context across conversational turns, align heterogeneous data representations, and generate appropriate multimodal responses, facilitating more versatile and natural communication.