Built independently by an author, for readers. Read the story and support ChapterPal

keyword

modality imbalance

Modality imbalance is a condition in multimodal machine learning where disparities in data distribution, representation quality, or learning dynamics cause one modality to dominate the training process over others. When models process inputs from diverse channels, such as text, images, or audio, differences in convergence speeds or information richness often lead the network to disproportionately rely on the dominant modality while suppressing feature learning in weaker or underrepresented ones. Consequently, this imbalance prevents the system from effectively exploiting complementary cross-modal relationships, resulting in suboptimal joint representations and degraded performance across varied testing conditions.

2 items

Learning Memory-Augmented Unidirectional Metrics for Cross-modality Person Re-identification

Learning Memory-Augmented Unidirectional Metrics for Cross-modality Person Re-identification

Jialun Liu, Yifan Sun, Feng Zhu, Hongbin Pei, Yi Yang, Wenhui Li

OrganizationsBaiduJilin UniversityXi'an Jiaotong UniversityZhejiang University

Why you should read this

Proposes a memory-augmented unidirectional metric learning method that suppresses cross-modality discrepancy by using modality-specific proxies and memory banks to associate features across modalities while effectively handling data imbalance.

This paper tackles the cross-modality person re-identification (re-ID) problem by suppressing the modality discrepancy. In cross-modality re-ID, the query and gallery images are in different modalities. Given a training identity, the popular deep classification baseline shares the same proxy (i.e., a weight vector in the last classification layer) for two modalities. We find that it has considerable tolerance for the modality gap, because the shared proxy acts as an intermediate relay between two modalities. In response, we propose a Memory-Augmented Unidirectional Metric (MAUM) learning method consisting of two novel designs, i.e., unidirectional metrics, and memory-based augmentation. Specifically, MAUM first learns modality-specific proxies (MS-Proxies) independently under each modality. Afterward, MAUM uses the already-learned MS-Proxies as the static references for pulling close the features in the counterpart modality. These two unidirectional metrics (IR image to RGB proxy and RGB image to IR proxy) jointly alleviate the relay effect and benefit cross-modality association. The cross-modality association is further enhanced by storing the MS-Proxies into memory banks to increase the reference diversity. Importantly, we show that MAUM improves cross-modality re-ID under the modality-balanced setting and gains extra robustness against the modality-imbalance problem. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate the superiority of MAUM over the state-of-the-art. The code will be available.

Added

2026-10-06

ReconBoost: Boosting Can Achieve Modality Reconcilement

ReconBoost: Boosting Can Achieve Modality Reconcilement

Cong Hua, Qianqian Xu, Shilong Bao, Zhiyong Yang, Qingming Huang

OrganizationsChinese Academy of SciencesInstitute of Computing Technology, Chinese Academy of SciencesInstitute of Information Engineering, Chinese Academy of SciencesUniversity of Chinese Academy of Sciences

Why you should read this

Proposes a gradient-boosting-inspired alternating learning framework called ReconBoost that mitigates modality competition by dynamically updating individual modalities sequentially with regularization to reconcile uni-modal exploitation and cross-modal fusion.

This paper explores a novel multi-modal alternating learning paradigm pursuing a reconciliation between the exploitation of uni-modal features and the exploration of cross-modal interactions. This is motivated by the fact that current paradigms of multi-modal learning tend to explore multi-modal features simultaneously. The resulting gradient prohibits further exploitation of the features in the weak modality, leading to modality competition, where the dominant modality overpowers the learning process. To address this issue, we study the modality-alternating learning paradigm to achieve reconcilement. Specifically, we propose a new method called ReconBoost to update a fixed modality each time. Herein, the learning objective is dynamically adjusted with a reconcilement regularization against competition with the historical models. By choosing a KL-based reconcilement, we show that the proposed method resembles Friedman’s Gradient-Boosting (GB) algorithm, where the updated learner can correct errors made by others and help enhance the overall performance. The major difference with the classic GB is that we only preserve the newest model for each modality to avoid overfitting caused by ensembling strong learners. Furthermore, we propose a memory consolidation scheme and a global rectification scheme to make this strategy more effective. Experiments over six multi-modal benchmarks speak to the efficacy of the method. We release the code at https://github.com/huacong/ReconBoost.

Added

2026-09-26