keyword
multimodal classification
Multimodal classification is a machine learning task that assigns data to predefined categories by integrating and analyzing information from multiple distinct modalities, such as text, images, audio, video, or sensor signals. Unlike unimodal approaches that rely on a single type of input, multimodal classification extracts and aligns complementary features from diverse sources through various data fusion strategies. By synthesizing these heterogeneous representations, models are able to capture complex cross-modal relationships, achieve higher predictive accuracy, and maintain robust performance even when individual modalities are noisy, degraded, or incomplete.
2 items

Calibrating Multimodal Learning
Huan Ma, Qingyang Zhang, Changqing Zhang, Bingzhe Wu, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu
Why you should read this
Proposes a lightweight regularization method that prevents multimodal classifiers from becoming spuriously more confident when modalities are removed or corrupted, ensuring trustworthy uncertainty estimation across diverse architectures.
Multimodal machine learning has achieved remarkable progress in a wide range of scenarios. However, the reliability of multimodal learning remains largely unexplored. In this paper, through extensive empirical studies, we identify current multimodal classification methods suffer from unreliable predictive confidence that tend to rely on partial modalities when estimating confidence. Specifically, we find that the confidence estimated by current models could even increase when some modalities are corrupted. To address the issue, we introduce an intuitive principle for multimodal learning, i.e., the confidence should not increase when one modality is removed. Accordingly, we propose a novel regularization technique, i.e., Calibrating Multimodal Learning (CML) regularization, to calibrate the predictive confidence of previous methods. This technique could be flexibly equipped by existing models and improve the performance in terms of confidence calibration, classification accuracy, and model robustness.
Added
2026-10-05

Provable Dynamic Fusion for Low-Quality Multimodal Data
Qingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu, Huazhu Fu, Joey Tianyi Zhou, Xi Peng
Why you should read this
Establishes theoretical generalization bounds for dynamic decision-level multimodal integration and introduces Quality-aware Multimodal Fusion to prevent performance degradation on noisy and low-quality inputs.
The inherent challenge of multimodal fusion is to precisely capture the cross-modal correlation and flexibly conduct cross-modal interaction. To fully release the value of each modality and mitigate the influence of low-quality multimodal data, dynamic multimodal fusion emerges as a promising learning paradigm. Despite its widespread use, theoretical justifications in this field are still notably lacking. Can we design a provably robust multimodal fusion method? This paper provides theoretical understandings to answer this question under a most popular multimodal fusion framework from the generalization perspective. We proceed to reveal that several uncertainty estimation solutions are naturally available to achieve robust multimodal fusion. Then a novel multimodal fusion framework termed Quality-aware Multimodal Fusion (QMF) is proposed, which can improve the performance in terms of classification accuracy and model robustness. Extensive experimental results on multiple benchmarks can support our findings.
Added
2026-09-26
