Self-supervised vision representation learning is a machine learning approach in computer vision that trains neural networks to extract meaningful visual features from unlabeled images or videos without relying on human annotations. Instead of using manual labels, this paradigm generates supervisory signals directly from the inherent structure of the visual data through pretext tasks such as masked image modeling, contrastive learning, and instance discrimination across transformed views. By learning to predict missing visual regions, reconstruct content, or align representations across different augmentations of the same input, models capture both high-level semantic relationships and fine-grained spatial structures. The resulting representations provide a general-purpose feature foundation that can be transferred or fine-tuned for a wide variety of downstream vision tasks, including image classification, object detection, and semantic segmentation.