Built independently by an author, for readers. Read the story and support ChapterPal

keyword

audio-visual model

An audio-visual model is a multimodal computational system designed to simultaneously process, analyze, and integrate information from both auditory and visual data streams. By exploiting the complementary correlations between sound and imagery, such models learn unified cross-modal representations that overcome the limitations of unimodal processing. These architectures typically utilize specialized neural encoders to extract features from audio signals and video frames, which are then combined using multimodal fusion mechanisms, joint embedding spaces, or cross-attention networks. This integrated approach enables the system to capture temporal and spatial synchronization between what is heard and seen, facilitating tasks such as audio-visual speech recognition, sound source localization, active speaker detection, action classification, and multimedia generation.

1 item

PMR: Prototypical Modal Rebalance for Multimodal Learning

PMR: Prototypical Modal Rebalance for Multimodal Learning

Yunfeng Fan, Wenchao Xu, Haozhao Wang, Junxiao Wang, Song Guo

OrganizationsHong Kong Polytechnic UniversityHuazhong University of Science and TechnologyKing Abdullah University of Science and TechnologySDAIA-KAUST AI Center

Why you should read this

Proposes a model-agnostic strategy called Prototypical Modal Rebalance that uses class prototypes to actively guide the representation learning of underperforming modalities while preventing dominant modalities from prematurely converging.

Multimodal learning (MML) aims to jointly exploit the common priors of different modalities to compensate for their inherent limitations. However, existing MML methods often optimize a uniform objective for different modalities, leading to the notorious “modality imbalance” problem and counterproductive MML performance. To address the problem, some existing methods modulate the learning pace based on the fused modality, which is dominated by the better modality and eventually results in a limited improvement on the worse modal. To better exploit the features of multimodal, we propose Prototypical Modality Rebalance (PMR) to perform stimulation on the particular slow-learning modality without interference from other modalities. Specifically, we introduce the prototypes that represent general features for each class, to build the non-parametric classifiers for uni-modal performance evaluation. Then, we try to accelerate the slow-learning modality by enhancing its clustering toward prototypes. Furthermore, to alleviate the suppression from the dominant modality, we introduce a prototype-based entropy regularization term during the early training stage to prevent premature convergence. Besides, our method only relies on the representations of each modality and without restrictions from model structures and fusion methods, making it with great application potential for various scenarios. The source code is available here¹.

Added

2026-09-26