ReCo: Retrieve and Co-segment for Zero-shot Transfer
Gyungin ShinWeidi XieSamuel Albanie
Proposes a zero-shot semantic segmentation framework that combines vision-language image retrieval with cross-image co-segmentation to build open-vocabulary segmenters from unlabeled data without requiring any manual pixel annotations.
Semantic segmentation—identifying and outlining specific concepts within images at the pixel level—is critical for domains such as autonomous driving, healthcare, and industrial inspection. However, practical deployment is bottlenecked by the prohibitive expense of collecting manual pixel annotations, which can take up to 90 minutes per image. While unsupervised methods eliminate manual labeling costs, they typically require labeled target examples simply to assign names to predicted regions, and they struggle to identify rare or novel concepts described in open text.
The article demonstrates an open-vocabulary segmentation framework, termed Retrieve and Co-segment (ReCo), that enables zero-shot semantic segmentation without requiring pixel-level supervision or labeled examples from the target domain. The approach evaluates whether combining pre-trained vision-language models with deep visual correspondences can dynamically generate accurate segmenters for arbitrary categories on the fly.
The methodology operates in three main stages without downstream model training. First, the pre-trained CLIP model queries a large unlabelled image repository (such as ImageNet or billions of web images) using text prompts to retrieve relevant candidate images. Second, a vision encoder finds recurring seed pixels across these retrieved images to isolate the target concept via co-segmentation, assisted by language gating and context elimination to filter out background distractors like sky or roads. Finally, the resulting reference embeddings and dense visual-language saliency maps are applied directly to segment target images. When target data is accessible, an optional extension, ReCo+, trains a standard segmentation network on ReCo's initial predictions.
The key findings show substantial performance improvements across multiple standard benchmarks. In zero-shot transfer settings, ReCo achieved 27.2% mean intersection-over-union (mIoU) on COCO-Stuff compared to 19.8% for prior models, reached 22.0% mIoU on Cityscapes versus 10.0% for DenseCLIP, and scored 29.8% mIoU on KITTI-STEP compared to 15.3% for previous approaches. When employing unsupervised adaptation (ReCo+), the framework outperformed existing baselines on Cityscapes with 83.7% pixel accuracy and 24.2% mIoU, and achieved 31.9% mIoU on KITTI-STEP. Furthermore, the framework demonstrated a rare ability to segment novel and specialized categories without training labels, achieving 93.3% accuracy and 44.9% IoU on rare fire extinguisher instances and successfully isolating unique objects such as the Antikythera mechanism.
These results demonstrate that organizations can bypass costly manual pixel annotation pipelines and deploy flexible, open-vocabulary image segmentation models directly from text queries. This dramatically lowers deployment timelines, cost, and complexity when expanding systems to novel environments or rare objects. If target imagery is available, teams can choose the ReCo+ adaptation trade-off to boost accuracy further, while baseline ReCo provides immediate zero-shot deployment.
Decision-makers should consider ReCo as an effective proof-of-concept for rapid semantic discovery and exploratory segmentation pipelines. However, organizations should proceed cautiously before operational deployment. The approach relies heavily on large-scale foundation models that require significant computational resources, contains subtle optimization choices guided by standard benchmarks, and depends on uncurated web datasets that may inherit demographic biases or lack ultra-specific concepts. Next steps should focus on distilling the visual and language models into lightweight architectures and implementing rigorous moderation and data-governance safeguards on the retrieval archives.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. It introduces the foundational CLIP architecture and contrastive vision-language pretraining upon which ReCo directly relies for text-driven image retrieval and zero-shot transfer.
- Paper: Unsupervised Semantic Segmentation by Distilling Feature Correspondences, Mark Hamilton et al. (2022). It establishes unsupervised semantic segmentation via deep feature correspondences across images, providing the visual correspondence principles adapted by ReCo's co-segmentation stage.
- Paper: Open-vocabulary Object Detection via Vision and Language Knowledge Distillation, Xiuye Gu et al. (2021). It pioneers open-vocabulary localization by distilling pretrained vision-language representations, framing the problem setting that ReCo advances to pixel-level zero-shot segmentation.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). It introduces end-to-end fully convolutional pixel-level classification, serving as the architectural baseline for dense semantic segmentation models fine-tuned in ReCo+.
- Paper: Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning, Jishnu Mukhoti et al. (2023). It builds directly upon zero-shot open-vocabulary segmentation by learning patch-aligned contrastive representations from image-text pairs without explicit retrieval and co-segmentation pipelines.
- Paper: Hierarchical Open-vocabulary Universal Image Segmentation, Xudong Wang et al. (2023). It extends open-vocabulary universal image segmentation into a unified multi-granularity framework that handles both foreground things and background stuff hierarchically.
- Paper: Segment Anything, Alexander M. Kirillov et al. (2023). It advances zero-shot segmentation to a foundational scale by introducing promptable mask generation trained on massive cross-domain visual data.
- Paper: SAM 3: Segment Anything with Concepts, Nicolas Carion et al. (2025). It unifies promptable concept detection and open-vocabulary segmentation across images and video, representing a major subsequent milestone in zero-shot concept localization.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). It provides a comprehensive survey and taxonomy of vision-language models applied across dense visual downstream tasks, contextualizing retrieval and zero-shot transfer methods like ReCo.
