The Retrieve and Co-segment framework is a computer vision approach designed to perform semantic image segmentation without requiring manually annotated pixel-level labels. It functions by utilizing a pre-trained vision-language model to retrieve relevant, unlabeled images matching specific textual concept names from large image pools. The framework then employs co-segmentation techniques across the retrieved image groups to automatically detect and extract recurring visual entities, producing synthetic segmentation masks. These pseudo-labeled samples are subsequently used to train a segmentation model capable of zero-shot transfer across large vocabularies, enabling precise object localization and classification for standard as well as rare visual concepts.