keyword
cross-view prediction
Cross-view prediction is a self-supervised representation learning technique in which a machine learning model is trained to predict, reconstruct, or infer one view, channel, or modality of a data sample given another view of the same underlying entity or scene. In this framework, different perspectives of an input—such as separate color channels, audio and visual streams, distinct sensor readings, or spatial viewpoints—serve mutually as inputs and supervision targets. By learning to map between complementary yet incomplete observations without human annotations, the model captures shared structural, geometric, and semantic factors that remain consistent across varying observations. Unlike contrastive methods that optimize representation distances through instance discrimination in an embedding space, cross-view prediction typically relies on generative or predictive reconstruction objectives across different perceptual domains.
1 item

