keyword
multi-modal classification
Multi-modal classification is a machine learning task that categorizes data instances by jointly analyzing and integrating information from two or more distinct input types, such as text, images, audio, video, or sensor signals. Unlike unimodal classification, which relies on a single data stream, multi-modal classification leverages the complementary and correlated characteristics of different modalities to improve overall predictive accuracy and robustness. Systems designed for this task extract feature representations from each individual modality and combine them using various fusion strategies, such as early, intermediate, or late fusion, allowing the model to capture complex cross-modal relationships and output a unified class prediction.
1 item

