Unsupervised dense prediction tasks are computer vision objectives that require generating fine-grained, pixel-level predictions across an entire image without relying on human-annotated labels or ground-truth masks during training. Unlike global tasks that output a single prediction for a whole image, dense prediction generates localized outputs for every individual pixel or spatial coordinate, such as category assignments, continuous depth values, or motion vectors. Common examples include unsupervised semantic segmentation, unsupervised monocular depth estimation, optical flow estimation, and salient object discovery. To solve these tasks without manual annotations, models typically exploit self-supervised visual representations, spatial and geometric constraints, feature similarity, or clustering mechanisms to automatically discover visual structures and group semantically or physically consistent regions directly from unlabelled data.