keyword
cross-modal pseudo-labeling framework
A cross-modal pseudo-labeling framework is a machine learning system that automatically generates synthetic training annotations for one data modality by leveraging semantic information transferred from a complementary modality. In multimodal tasks such as open-vocabulary visual perception, the framework aligns shared representations across distinct data types, such as associating textual descriptions with corresponding visual features or spatial regions. These derived alignments produce pseudo-labels, including bounding boxes, segmentation masks, or category tags, without requiring exhaustive human annotations. The resulting pseudo-labels are then used to supervise or self-train downstream models on novel concepts, frequently utilizing filtering, confidence estimation, or distillation techniques to mitigate the adverse effects of label noise inherent in the automated generation process.
1 item

