Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multi-modal fusion methods

Multi-modal fusion methods are machine learning and data processing techniques that integrate information from two or more distinct data modalities, such as text, images, audio, video, or sensor streams, into a single coherent framework to make predictions or perform analytical tasks. These approaches aim to capture both modality-specific features and cross-modal interactions, allowing an automated system to leverage complementary information and achieve greater accuracy than unimodal models. Fusion strategies are commonly categorized by the stage at which data is combined, including early fusion, which joins raw or low-level features at the input stage; late fusion, which aggregates the final predictions of separate modality-specific models; and intermediate or joint fusion, which merges learned representations within hidden layers of deep neural networks using mechanisms such as cross-attention, tensor operations, or shared latent spaces. By synthesizing diverse information streams, multi-modal fusion methods enable models to handle complex, heterogeneous environments while addressing practical challenges such as modality imbalance, noise, and missing data.

1 item

ReconBoost: Boosting Can Achieve Modality Reconcilement

ReconBoost: Boosting Can Achieve Modality Reconcilement

Cong Hua, Qianqian Xu, Shilong Bao, Zhiyong Yang, Qingming Huang

OrganizationsChinese Academy of SciencesInstitute of Computing Technology, Chinese Academy of SciencesInstitute of Information Engineering, Chinese Academy of SciencesUniversity of Chinese Academy of Sciences

Why you should read this

Proposes a gradient-boosting-inspired alternating learning framework called ReconBoost that mitigates modality competition by dynamically updating individual modalities sequentially with regularization to reconcile uni-modal exploitation and cross-modal fusion.

This paper explores a novel multi-modal alternating learning paradigm pursuing a reconciliation between the exploitation of uni-modal features and the exploration of cross-modal interactions. This is motivated by the fact that current paradigms of multi-modal learning tend to explore multi-modal features simultaneously. The resulting gradient prohibits further exploitation of the features in the weak modality, leading to modality competition, where the dominant modality overpowers the learning process. To address this issue, we study the modality-alternating learning paradigm to achieve reconcilement. Specifically, we propose a new method called ReconBoost to update a fixed modality each time. Herein, the learning objective is dynamically adjusted with a reconcilement regularization against competition with the historical models. By choosing a KL-based reconcilement, we show that the proposed method resembles Friedman’s Gradient-Boosting (GB) algorithm, where the updated learner can correct errors made by others and help enhance the overall performance. The major difference with the classic GB is that we only preserve the newest model for each modality to avoid overfitting caused by ensembling strong learners. Furthermore, we propose a memory consolidation scheme and a global rectification scheme to make this strategy more effective. Experiments over six multi-modal benchmarks speak to the efficacy of the method. We release the code at https://github.com/huacong/ReconBoost.

Added

2026-09-26