keyword
self-supervised ViT
A self-supervised Vision Transformer (ViT) is a deep learning model based on the transformer architecture that learns visual representations from unlabeled image data without relying on human annotations. Instead of using predefined class labels, it is pre-trained through self-supervised objectives such as contrastive learning, self-distillation, or masked patch reconstruction, which encourage the network to discover visual patterns and underlying structures autonomously. By dividing images into discrete patches and processing them with multi-head self-attention mechanisms, a self-supervised ViT captures both global context and fine-grained spatial relationships, often naturally producing attention maps that delineate object boundaries and group semantically consistent regions. Consequently, these models generate versatile feature representations that serve as robust foundations for various downstream computer vision tasks, including image classification, object detection, and dense unsupervised semantic segmentation.
1 item

