keyword
optimal multimodal augmentation
An optimal multimodal augmentation refers to a transformed version of multimodal data where the only information preserved between the original inputs and the augmented counterparts is strictly relevant to a downstream task, free from task-irrelevant noise. Formulated within the framework of self-supervised representation learning and information theory, this condition occurs when the mutual information between the original multimodal inputs and their augmented views exactly equals the mutual information between the original inputs and the target task label. In practice, such augmentations perfectly preserve all essential task signals, capturing both redundant features shared across modalities and distinct features unique to individual modalities, while altering or removing superfluous background details and noise. By maintaining task-relevant properties invariant across views, optimal multimodal augmentations provide a theoretical benchmark that allows models to isolate sufficient, task-specific representations from unlabeled multimodal data.
1 item

