Built independently by an author, for readers. Read the story and support ChapterPal

keyword

object co-segmentation

Object co-segmentation is a computer vision task that involves jointly identifying and segmenting common foreground objects across a collection of two or more images. Unlike conventional image segmentation methods that analyze each image in isolation, co-segmentation exploits visual correspondences and shared semantic features across multiple images to separate recurring subjects from varying backgrounds. This approach allows systems to delineate either the same physical object or different instances belonging to the same semantic category, often operating in unsupervised or weakly supervised settings without requiring manual pixel-level annotations. As a result, it is widely utilized for tasks such as automated object discovery, image retrieval, multi-view 3D reconstruction, and generating labeled visual data for training downstream models.

1 item

ReCo: Retrieve and Co-segment for Zero-shot Transfer

ReCo: Retrieve and Co-segment for Zero-shot Transfer

Gyungin Shin, Weidi Xie, Samuel Albanie

OrganizationsDepartment of EngineeringShanghai Jiao Tong UniversityUniversity of CambridgeUniversity of Oxford

Why you should read this

Proposes a zero-shot semantic segmentation framework that combines vision-language image retrieval with cross-image co-segmentation to build open-vocabulary segmenters from unlabeled data without requiring any manual pixel annotations.

Semantic segmentation has a broad range of applications, but its real-world impact has been significantly limited by the prohibitive annotation costs necessary to enable deployment. Segmentation methods that forgo supervision can side-step these costs, but exhibit the inconvenient requirement to provide labelled examples from the target distribution to assign concept names to predictions. An alternative line of work in language-image pre-training has recently demonstrated the potential to produce models that can both assign names across large vocabularies of concepts and enable zero-shot transfer for classification, but do not demonstrate commensurate segmentation abilities. We leverage the retrieval abilities of one such language-image pre-trained model, CLIP, to dynamically curate training sets from unlabelled images for arbitrary collections of concept names, and leverage the robust correspondences offered by modern image representations to co-segment entities among the resulting collections. The synthetic segment collections are then employed to construct a segmentation model (without requiring pixel labels) whose knowledge of concepts is inherited from the scalable pre-training process of CLIP. We demonstrate that our approach, termed Retrieve and Co-segment (ReCo) performs favourably to conventional unsupervised segmentation approaches while inheriting the convenience of nameable predictions and zero-shot transfer. We also demonstrate ReCo’s ability to generate specialist segmenters for extremely rare objects.

Added

2026-09-26