ReCo: Retrieve and Co-segment for Zero-shot Transfer

Gyungin ShinWeidi XieSamuel Albanie

article2022NeurIPS142 citations

Proposes a zero-shot semantic segmentation framework that combines vision-language image retrieval with cross-image co-segmentation to build open-vocabulary segmenters from unlabeled data without requiring any manual pixel annotations.

Listen

Semantic segmentation—identifying and outlining specific concepts within images at the pixel level—is critical for domains such as autonomous driving, healthcare, and industrial inspection. However, practical deployment is bottlenecked by the prohibitive expense of collecting manual pixel annotations, which can take up to 90 minutes per image. While unsupervised methods eliminate manual labeling costs, they typically require labeled target examples simply to assign names to predicted regions, and they struggle to identify rare or novel concepts described in open text.

The article demonstrates an open-vocabulary segmentation framework, termed Retrieve and Co-segment (ReCo), that enables zero-shot semantic segmentation without requiring pixel-level supervision or labeled examples from the target domain. The approach evaluates whether combining pre-trained vision-language models with deep visual correspondences can dynamically generate accurate segmenters for arbitrary categories on the fly.

The methodology operates in three main stages without downstream model training. First, the pre-trained CLIP model queries a large unlabelled image repository (such as ImageNet or billions of web images) using text prompts to retrieve relevant candidate images. Second, a vision encoder finds recurring seed pixels across these retrieved images to isolate the target concept via co-segmentation, assisted by language gating and context elimination to filter out background distractors like sky or roads. Finally, the resulting reference embeddings and dense visual-language saliency maps are applied directly to segment target images. When target data is accessible, an optional extension, ReCo+, trains a standard segmentation network on ReCo's initial predictions.

The key findings show substantial performance improvements across multiple standard benchmarks. In zero-shot transfer settings, ReCo achieved 27.2% mean intersection-over-union (mIoU) on COCO-Stuff compared to 19.8% for prior models, reached 22.0% mIoU on Cityscapes versus 10.0% for DenseCLIP, and scored 29.8% mIoU on KITTI-STEP compared to 15.3% for previous approaches. When employing unsupervised adaptation (ReCo+), the framework outperformed existing baselines on Cityscapes with 83.7% pixel accuracy and 24.2% mIoU, and achieved 31.9% mIoU on KITTI-STEP. Furthermore, the framework demonstrated a rare ability to segment novel and specialized categories without training labels, achieving 93.3% accuracy and 44.9% IoU on rare fire extinguisher instances and successfully isolating unique objects such as the Antikythera mechanism.

These results demonstrate that organizations can bypass costly manual pixel annotation pipelines and deploy flexible, open-vocabulary image segmentation models directly from text queries. This dramatically lowers deployment timelines, cost, and complexity when expanding systems to novel environments or rare objects. If target imagery is available, teams can choose the ReCo+ adaptation trade-off to boost accuracy further, while baseline ReCo provides immediate zero-shot deployment.

Decision-makers should consider ReCo as an effective proof-of-concept for rapid semantic discovery and exploratory segmentation pipelines. However, organizations should proceed cautiously before operational deployment. The approach relies heavily on large-scale foundation models that require significant computational resources, contains subtle optimization choices guided by standard benchmarks, and depends on uncurated web datasets that may inherit demographic biases or lack ultra-specific concepts. Next steps should focus on distilling the visual and language models into lightweight architectures and implementing rigorous moderation and data-governance safeguards on the retrieval archives.

arXiv: 2206.07045
  • Paper: Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning, Jishnu Mukhoti et al. (2023). It builds directly upon zero-shot open-vocabulary segmentation by learning patch-aligned contrastive representations from image-text pairs without explicit retrieval and co-segmentation pipelines.
  • Paper: Hierarchical Open-vocabulary Universal Image Segmentation, Xudong Wang et al. (2023). It extends open-vocabulary universal image segmentation into a unified multi-granularity framework that handles both foreground things and background stuff hierarchically.
  • Paper: Segment Anything, Alexander M. Kirillov et al. (2023). It advances zero-shot segmentation to a foundational scale by introducing promptable mask generation trained on massive cross-domain visual data.
  • Paper: SAM 3: Segment Anything with Concepts, Nicolas Carion et al. (2025). It unifies promptable concept detection and open-vocabulary segmentation across images and video, representing a major subsequent milestone in zero-shot concept localization.
  • Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). It provides a comprehensive survey and taxonomy of vision-language models applied across dense visual downstream tasks, contextualizing retrieval and zero-shot transfer methods like ReCo.
Cover for ReCo: Retrieve and Co-segment for Zero-shot Transfer

Abstract

Semantic segmentation has a broad range of applications, but its real-world impact has been significantly limited by the prohibitive annotation costs necessary to enable deployment. Segmentation methods that forgo supervision can side-step these costs, but exhibit the inconvenient requirement to provide labelled examples from the target distribution to assign concept names to predictions. An alternative line of work in language-image pre-training has recently demonstrated the potential to produce models that can both assign names across large vocabularies of concepts and enable zero-shot transfer for classification, but do not demonstrate commensurate segmentation abilities.

We leverage the retrieval abilities of one such language-image pre-trained model, CLIP, to dynamically curate training sets from unlabelled images for arbitrary collections of concept names, and leverage the robust correspondences offered by modern image representations to co-segment entities among the resulting collections. The synthetic segment collections are then employed to construct a segmentation model (without requiring pixel labels) whose knowledge of concepts is inherited from the scalable pre-training process of CLIP. We demonstrate that our approach, termed Retrieve and Co-segment (ReCo) performs favourably to conventional unsupervised segmentation approaches while inheriting the convenience of nameable predictions and zero-shot transfer. We also demonstrate ReCo’s ability to generate specialist segmenters for extremely rare objects.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Method
  • 3.1 Retrieve and Co-segment (ReCo)
  • 3.2 Language-guided co-segmentation and context elimination
  • 3.3 ReCo+: Fine-tuning ReCo with pseudo-labels
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Implementation details
  • 4.3 Ablation studies
  • 4.4 Comparison to state-of-the-art unsupervised methods
  • 4.5 Segmenting rare concepts
  • 4.6 Limitations
  • 5 Conclusion
  • 6 Acknowledgements
  • References
  • A Discussion of broader impact, supervision and data
  • A.1 Broader Impact
  • A.2 Supervisory signals for ReCo and prior work
  • A.3 Discussion of consent in used datasets
  • A.4 Discussion on whether data contains personally identifiable information or offensive content
  • A.5 Dataset licenses
  • B Experiment details
  • B.1 Pseudocode for ReCo
  • B.2 Prompt engineering
  • B.3 Details of ablation study to assess CLIP retrieval performance
  • B.4 Hyperparameters for ReCo+ training
  • C Additional ablation studies
  • C.1 Choices of context categories
  • C.2 Category name rephrasing to reduce ambiguity
  • D Additional visualisations

Knowls

  1. Knowl 1 — Unsupervised Semantic Segmentation with Language-Image Pre-training

    definition

    Unsupervised semantic segmentation with language-image pre-training (USLIP) is the task of assigning to each pixel location ω∈Ω\omega \in \Omega of an input image x∈R3×h×wx \in \mathbb{R}^{3 \times h \times w} (with spatial domain Ω={1,…,h−1}×{1,…,w−1}\Omega = \{1, \dots, h-1\} \times \{1, \dots, w-1\}) a semantic label c∈Cc \in \mathcal{C} from a set of ∣C∣|\mathcal{C}| mutually exclusive target categories, without relying on human pixel-level annotations or image-level downstream task supervision.

    USLIP assumes access to a large pre-training corpus of paired image and text data (e.g., via models like CLIP). Unlike conventional unsupervised semantic segmentation, which outputs unlabelled clusters requiring post-hoc category assignment via Hungarian matching or nearest-neighbor search on ground-truth target annotations, USLIP produces directly nameable category predictions and performs zero-shot transfer on target datasets without access to the target data distribution during training.

  2. Knowl 2 — Retrieve and Co-segment (ReCo) Framework

    model/method

    Retrieve and Co-segment (ReCo) is a zero-shot semantic segmentation framework that generates segmentations for arbitrary concept names on the fly without fine-tuning or access to target-domain images. The framework operates in three sequential stages:

    1. Exemplar Retrieval: Given a list of target concepts C\mathcal{C} and an unlabelled image collection U\mathcal{U}, ReCo uses the CLIP text encoder ψT\psi_T to generate a text query embedding ψT(c)∈Re\psi_T(c) \in \mathbb{R}^e for each category c∈Cc \in \mathcal{C}. It retrieves an archive of kk nearest-neighbor images from U\mathcal{U} based on cosine similarity with CLIP image embeddings ψI(x)∈Re\psi_I(x) \in \mathbb{R}^e.

    2. Seed-Pixel Co-segmentation: For the kk retrieved archive images of category cc, dense visual features are extracted using a pre-trained image encoder ϕI\phi_I (such as DeiT-S/16-SIN). By computing cross-image correspondences, ReCo discovers the most salient and consistent pixel (seed pixel) in each archive image and averages their features to yield an L2-normalized reference concept embedding fc∈Rdf_c \in \mathbb{R}^d.

    3. Inference with Multi-Modal Refinement: For an unseen test image xnewx_{\text{new}}, an initial probability map is obtained by computing the dot product between fcf_c and the dense visual features of xnewx_{\text{new}}, which is subsequently modulated by a DenseCLIP text-visual saliency map to produce the final per-pixel category prediction.

  3. Knowl 3 — Seed Pixel Extraction and Reference Embedding Computation

    algorithm

    Given an archive of kk retrieved images for category cc and a dense visual encoder ϕI\phi_I, seed pixel discovery identifies the most representative pixel embedding across the archive to build a single reference embedding fc∈Rdf_c \in \mathbb{R}^d.

    Input: Archive images {x1,…,xk}\{x_1, \dots, x_k\}, visual feature extractor ϕI\phi_I
    Output: L2-normalized reference concept embedding fc∈Rdf_c \in \mathbb{R}^d
    for i=1i = 1 to kk do
        Extract dense feature map Fi=ϕI(xi)∈Rd×h×wF_i = \phi_I(x_i) \in \mathbb{R}^{d \times h \times w}
        Normalize FiF_i along its channel dimension: Fi←Fi/∥Fi∥2F_i \leftarrow F_i / \|F_i\|_2
    end for
    Construct adjacency matrix blocks Ai,j=Fi⊤Fj∈Rhw×hwA_{i,j} = F_i^\top F_j \in \mathbb{R}^{hw \times hw} for all i,j∈{1,…,k}i, j \in \{1, \dots, k\}
    for i=1i = 1 to kk do
        for j=1j = 1 to kk do
            Compute column-wise maximum: Mi,j=max⁡cols(Ai,j)∈Rhw×1M_{i,j} = \max_{\text{cols}}(A_{i,j}) \in \mathbb{R}^{hw \times 1}
        end for
        Compute mean support across all kk images: Si=1k∑j=1kMi,j∈Rhw×1S_i = \frac{1}{k} \sum_{j=1}^k M_{i,j} \in \mathbb{R}^{hw \times 1}
        Find spatial index of maximum support: pi∗=argmax⁡(Si)∈{1,…,hw}p_i^* = \operatorname{argmax}(S_i) \in \{1, \dots, hw\}
        Extract seed pixel feature vector: si=Fi(pi∗)∈Rds_i = F_i(p_i^*) \in \mathbb{R}^d
    end for
    Compute average seed embedding and normalize: fc=∑i=1ksi∥∑i=1ksi∥2f_c = \frac{\sum_{i=1}^k s_i}{\|\sum_{i=1}^k s_i\|_2}
    return fcf_c
  4. Knowl 4 — DenseCLIP-Modulated Inference in ReCo

    equation

    For a test image xnewx_{\text{new}}, dense visual features Fnew∈Rd×h×wF_{\text{new}} \in \mathbb{R}^{d \times h \times w} are extracted using visual encoder ϕI\phi_I and L2-normalized along channels. The initial probability map Pnewc∈[0,1]h×wP_{\text{new}}^c \in [0, 1]^{h \times w} for category cc with reference embedding fc∈Rdf_c \in \mathbb{R}^d is computed as:

    Pnewc=σ(fc⋅Fnew)P_{\text{new}}^c = \sigma(f_c \cdot F_{\text{new}})

    where σ\sigma denotes the sigmoid activation function and ⋅\cdot is the spatial channel-wise dot product.

    To construct the text-driven saliency map Snewc∈[0,1]h×wS_{\text{new}}^c \in [0, 1]^{h \times w}, features Vnew∈Rev×h×wV_{\text{new}} \in \mathbb{R}^{e_v \times h \times w} are extracted from the values of the final self-attention layer of the CLIP image encoder, projected into the joint vision-language space Re\mathbb{R}^e via the CLIP image projection matrix, and L2-normalized. The L2-normalized CLIP text embedding ψT(c)∈Re\psi_T(c) \in \mathbb{R}^e is convolved as a 1×11 \times 1 filter over these projected features, followed by a sigmoid activation:

    Snewc=σ(proj⁡(Vnew)∗ψT(c))S_{\text{new}}^c = \sigma(\operatorname{proj}(V_{\text{new}}) * \psi_T(c))

    The final prediction map Pˉnewc∈[0,1]h×w\bar{P}_{\text{new}}^c \in [0, 1]^{h \times w} is defined by the Hadamard (element-wise) product:

    Pˉnewc=Pnewc⊙Snewc\bar{P}_{\text{new}}^c = P_{\text{new}}^c \odot S_{\text{new}}^c

    For multi-class segmentation, the class label at each pixel is assigned via argmax⁡c∈CPˉnewc\operatorname{argmax}_{c \in \mathcal{C}} \bar{P}_{\text{new}}^c, optionally followed by dense Conditional Random Field (CRF) post-processing.

  5. Knowl 5 — Language-Guided Co-segmentation and Context Elimination

    model/method

    To prevent the co-segmentation process from incorrectly selecting dominant co-occurring background elements (such as sky with aeroplanes, or roads with cars), ReCo modifies the pairwise similarity submatrices Aj,i∈Rhw×hwA_{j,i} \in \mathbb{R}^{hw \times hw} (encoding similarities from image jj to image ii) using two filtering mechanisms before seed pixel selection:

    1. Language-Guided Co-segmentation (LGC): For each archive image xix_i, a DenseCLIP saliency map Sic∈[0,1]h×wS_i^c \in [0, 1]^{h \times w} for target concept cc is computed and vectorised into vec⁡(Sic)∈[0,1]1×hw\operatorname{vec}(S_i^c) \in [0, 1]^{1 \times hw}. Each row of Aj,iA_{j,i} is replaced with its element-wise product with vec⁡(Sic)\operatorname{vec}(S_i^c) for all j∈{1,…,k}j \in \{1, \dots, k\}, gating out similarities for pixels not highlighted by CLIP.

    2. Context Elimination (CE): A set of common background distractor classes c~∈{tree,sky,building,road,person}\tilde{c} \in \{\text{tree}, \text{sky}, \text{building}, \text{road}, \text{person}\} is pre-defined, and their reference embeddings fc~f_{\tilde{c}} are pre-computed. The background probability map Pic~=σ(fc~⋅Fi)P_i^{\tilde{c}} = \sigma(f_{\tilde{c}} \cdot F_i) suppresses contextual background regions by replacing each row of Aj,iA_{j,i} with its element-wise product with vec⁡(1−Pic~)∈[0,1]1×hw\operatorname{vec}(1 - P_i^{\tilde{c}}) \in [0, 1]^{1 \times hw}. If target class cc is itself one of the background categories c~\tilde{c}, fcf_c is directly replaced with fc~f_{\tilde{c}}.

  6. Knowl 6 — ReCo+: Unsupervised Target Adaptation via Pseudo-Label Self-Training

    model/method

    When unlabelled images from the target distribution are accessible at training time, ReCo+ adapts to the target domain through self-training without using any manual ground-truth labels.

    ReCo first runs zero-shot inference on the unlabelled target dataset images to generate pseudo-segmentation masks. A standard supervised segmentation network (DeepLabv3+ with a ResNet-101 backbone) is then trained on these pseudo-masks. Training hyperparameters comprise:

    • Optimizer: Adam with initial learning rate 5×10−45 \times 10^{-4} and weight decay 2×10−42 \times 10^{-4}
    • Learning rate schedule: Polynomial decay schedule
    • Training duration: 20,000 iterations with batch size 8 (approx. 5 hours on an NVIDIA P40 GPU)
    • Data augmentations: Random scaling, random cropping to 320×320320 \times 320 pixels, horizontal flipping, random color jittering, and Gaussian blurring.
  7. Knowl 7 — Component Ablation of the ReCo Framework on PASCAL-Context

    data/table

    Ablation results on the PASCAL-Context validation set (59 categories) evaluate the individual and combined contributions of DenseCLIP inference modulation, Language-Guided Co-segmentation (LGC), Context Elimination (CE), and CRF post-processing. All setups utilize a DeiT-S/16-SIN visual backbone for dense co-segmentation and an archive of k=50k=50 images curated from ImageNet-1K using ViT-L/14@336px CLIP.

    DenseCLIP LGC CE CRF Pixel Acc. (%) mIoU (%)
    ✗ ✗ ✗ ✗ 16.8 5.7
    ✓ ✗ ✗ ✗ 41.1 21.8
    ✓ ✓ ✗ ✗ 43.1 23.1
    ✓ ✗ ✓ ✗ 49.7 26.0
    ✓ ✓ ✓ ✗ 50.9 26.6
    ✓ ✓ ✓ ✓ 51.6 27.2

    DenseCLIP provides the single largest gain (+16.1 mIoU), demonstrating the benefit of leveraging vision-language alignments directly at test time. Adding CE yields a +4.2 mIoU gain over base DenseCLIP, and combining LGC and CE achieves 26.6 mIoU (+4.8 over DenseCLIP alone). Final CRF refinement adds an additional +0.6 mIoU and +0.7% pixel accuracy.

  8. Knowl 8 — Comparative Benchmark Performance across COCO-Stuff, Cityscapes, and KITTI-STEP

    data/table

    Segmentation performance across three standard benchmarks comparing zero-shot transfer (no access to target data) and unsupervised adaptation (training on unlabelled target images). Evaluated using pixel accuracy (Acc.) and mean intersection-over-union (mIoU).

    COCO-Stuff Cityscapes KITTI-STEP
    Model Acc. (%) mIoU (%) Acc. (%) mIoU (%) Acc. (%) mIoU (%)
    Zero-shot transfer
    DenseCLIP 32.3 19.8 35.9 10.0 34.1 15.3
    ReCo (Ours) 46.6 27.2 65.4 22.0 70.6 29.8
    Unsupervised adaptation
    IIC 21.8 6.7 47.9 6.4 - -
    MDC 32.2 9.8 40.7 7.1 - -
    PiCIE 48.1 13.8 65.5 12.3 - -
    PiCIE + H 50.0 14.4 - - - -
    STEGO 56.9 28.2 73.2 21.0 - -
    SegSort - - - - 69.8 19.2
    HSG - - - - 73.8 21.7
    ReCo+ (Ours) 54.5 33.0 83.7 24.2 75.3 31.9

    Under zero-shot transfer, ReCo substantially outperforms DenseCLIP across all datasets (+7.4 mIoU on COCO-Stuff, +12.0 mIoU on Cityscapes, +14.5 mIoU on KITTI-STEP). When adapted via ReCo+, it achieves superior mIoU over prior unsupervised adaptation methods (including STEGO and HSG) on all three benchmarks.

  9. Knowl 9 — Zero-Shot Segmentation of Rare and Unannotated Concepts

    empirical result

    Using LAION-5B (a 5-billion unlabelled image dataset) as the underlying retrieval corpus, ReCo is able to construct specialist zero-shot segmenters for rare concepts that do not exist in standard segmentation benchmarks or WordNet.

    • Fire Extinguishers (FireNet): On 263 evaluation images from the FireNet dataset containing fire extinguishers (a category absent from ImageNet1K), ReCo achieves 93.3% pixel accuracy and 44.9% IoU without any training annotations.
    • Antikythera Mechanism: ReCo successfully retrieves unlabelled images and co-segments the ancient Greek mechanical computer (an entity absent from WordNet and labeled vision datasets), demonstrating open-world concept discovery and localization.
  10. Knowl 10 — Limitations of the ReCo Framework

    limitation

    The ReCo framework exhibits five primary limitations:

    1. Unrepresented Concepts: The retrieval assumption fails for hyper-specific or compositional concepts that do not exist in web-scale image datasets (e.g., 'a purple elephant with square orange feet wearing an inverted cowboy hat in front of the Doge's Palace').
    2. Inference Complexity: ReCo requires running dense forward passes through two distinct models during test time—a dense visual backbone (DeiT-S/16-SIN) and a vision-language model (CLIP ResNet50x16).
    3. Retrieval Corpus Bias: Using datasets like ImageNet1K introduces an object-centric bias into the reference embeddings.
    4. Static Pre-trained Vocabulary: Because vision-language models like CLIP are static and computationally expensive to retrain, ReCo cannot retrieve or segment concepts that emerged after the model's training date cutoff.
    5. Indirect Development Supervision: Hyperparameter selection and architectural design decisions were guided by ablation studies on the PASCAL-Context benchmark.

Coverage note — None was omitted. All contributed methods, algorithms, mathematical formulations, adaptations (ReCo+), ablation results, benchmark comparisons, rare-concept experiments, and limitations have been completely covered.

References

  1. 1.Jiwoon Ahn and Suha Kwak. Learning pixel-level semantic affinity with image-level supervision for weakly supervised semantic segmentation. In CVPR, 2018.
  2. 2.Sayan Banerjee, Avik Hati, Subhasis Chaudhuri, and Rajbabu Velmurugan. Cosegnet: Image co-segmentation using a conditional siamese convolutional network. In IJCAI, 2019.
  3. 3.Dhruv Batra, Adarsh Kowdle, Devi Parikh, Jiebo Luo, and Tsuhan Chen. icoseg: Interactive co-segmentation with intelligent scribble guidance. In CVPR, 2010.
  4. 4.Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei. What’s the point: Semantic segmentation with point supervision. In ECCV, 2016.
  5. 5.Maxime Bucher, Tuan-Hung Vu, Mathieu Cord, and Patrick Pérez. Zero-shot semantic segmentation. In NeurIPS, 2019.
  6. 6.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In CVPR, 2018.
  7. 7.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, 2021.
  8. 8.Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, and Ming-Hsuan Yang. Weakly-supervised semantic segmentation via sub-category exploration. In CVPR, 2020.
  9. 9.Hong Chen, Yifei Huang, and Hideki Nakayama. Semantic aware attention based deep object co-segmentation. In ACCV, 2018.
  10. 10.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018.
  11. 11.Bowen Cheng, Omkar Parkhi, and Alexander Kirillov. Pointly-supervised instance segmentation. arXiv:2104.06404, 2021.
  12. 12.Jang Hyun Cho, Utkarsh Mall, Kavita Bala, and Bharath Hariharan. Picie: Unsupervised semantic segmentation using invariance and equivariance in clustering. In CVPR, 2021.
  13. 13.Edo Collins, Radhakrishna Achanta, and Sabine Susstrunk. Deep feature factorization for concept discovery. In ECCV, 2018.
  14. 14.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016.
  15. 15.Jifeng Dai, Kaiming He, and Jian Sun. Boxsup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation. In ICCV, 2015.
  16. 16.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  17. 17.Jian Ding, Nan Xue, Gui-Song Xia, and Dengxin Dai. Decoupling zero-shot semantic segmentation. In CVPR, 2022.
  18. 18.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  19. 19.Ruochen Fan, Qibin Hou, Ming-Ming Cheng, Gang Yu, Ralph R. Martin, and Shi-Min Hu. Associating inter-image salient instances for weakly supervised semantic segmentation. In ECCV, 2018.
  20. 20.Robert Fergus, Li Fei-Fei, Pietro Perona, and Andrew Zisserman. Learning object categories from google’s image search. In ICCV, 2005.
  21. 21.Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin. Open-vocabulary image segmentation. arXiv:2112.12143, 2021.
  22. 22.Priya Goyal, Quentin Duval, Isaac Seessel, Mathilde Caron, Mannat Singh, Ishan Misra, Levent Sagun, Armand Joulin, and Piotr Bojanowski. Vision models are more robust and fair when pretrained on uncurated images without supervision. arXiv:2202.08360, 2022.
  23. 23.Zhangxuan Gu, Siyuan Zhou, Li Niu, Zihan Zhao, and Liqing Zhang. Context-aware feature generation for zero-shot semantic segmentation. In ACM MM, 2020.
  24. 24.Mark Hamilton, Zhoutong Zhang, Bharath Hariharan, Noah Snavely, and William T. Freeman. Unsupervised semantic segmentation by distilling feature correspondences. In ICLR, 2022.
  25. 25.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020.
  26. 26.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  27. 27.Yin-Yin He, Peizhen Zhang, Xiu-Shen Wei, Xiangyu Zhang, and Jian Sun. Relieving long-tailed instance segmentation via pairwise class balance. arXiv:2201.02784, 2022.
  28. 28.Kuang-Jui Hsu, Yen-Yu Lin, Yung-Yu Chuang, et al. Co-attention cnns for unsupervised object co-segmentation. In IJCAI, 2018.
  29. 29.Ping Hu, Stan Sclaroff, and Kate Saenko. Uncertainty-aware learning for zero-shot semantic segmentation. In NeurIPS, 2020.
  30. 30.Ronghang Hu, Piotr Dollár, Kaiming He, Trevor Darrell, and Ross Girshick. Learning to segment every thing. In CVPR, 2018.
  31. 31.Xinting Hu, Yi Jiang, Kaihua Tang, Jingyuan Chen, Chunyan Miao, and Hanwang Zhang. Learning to segment the tail. In CVPR, 2020.
  32. 32.Dat Huynh, Jason Kuen, Zhe Lin, Jiuxiang Gu, and Ehsan Elhamifar. Open-vocabulary instance segmentation via robust cross-modal pseudo-labeling. arXiv:2111.12698, 2021.
  33. 33.Jyh-Jing Hwang, Stella X Yu, Jianbo Shi, Maxwell D Collins, Tien-Ju Yang, Xiao Zhang, and Liang-Chieh Chen. Segsort: Segmentation by discriminative sorting of segments. In ICCV, 2019.
  34. 34.Xu Ji, João F Henriques, and Andrea Vedaldi. Invariant information clustering for unsupervised image classification and segmentation. In ICCV, 2019.
  35. 35.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, 2021.
  36. 36.Bin Jin, Maria V Ortiz Segovia, and Sabine Susstrunk. Webly supervised semantic segmentation. In CVPR, 2017.
  37. 37.Jeff Johnson, Matthijs Douze, and Hervé Jégou. Billion-scale similarity search with gpus. IEEE Transactions on Big Data, 2019.
  38. 38.Armand Joulin, Francis Bach, and Jean Ponce. Discriminative clustering for image co-segmentation. In CVPR, 2010.
  39. 39.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. Mdetr-modulated detection for end-to-end multi-modal understanding. In ICCV, 2021.
  40. 40.Naoki Kato, Toshihiko Yamasaki, and Kiyoharu Aizawa. Zero-shot semantic segmentation via variational mapping. In ICCVW, 2019.
  41. 41.Tsung-Wei Ke, Jyh-Jing Hwang, Yunhui Guo, Xudong Wang, and Stella X. Yu. Unsupervised hierarchical semantic segmentation with multiview cosegmentation and clustering transformers. In CVPR, 2022.
  42. 42.Anna Khoreva, Rodrigo Benenson, Jan Hosang, Matthias Hein, and Bernt Schiele. Simple does it: Weakly supervised instance and semantic segmentation. In CVPR, 2017.
  43. 43.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015.
  44. 44.Philipp Krähenbühl and Vladlen Koltun. Efficient inference in fully connected crfs with gaussian edge potentials. In NeurIPS, 2011.
  45. 45.Harold W. Kuhn. The Hungarian Method for the Assignment Problem. Naval Research Logistics Quarterly, 1955.
  46. 46.Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic segmentation. In ICLR, 2022.
  47. 47.Peike Li, Yunchao Wei, and Yi Yang. Consistent structural relation learning for zero-shot segmentation. In NeurIPS, 2020.
  48. 48.Weihao Li, Omid Hosseini Jafari, and Carsten Rother. Deep object co-segmentation. In ACCV, 2019.
  49. 49.Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. Scribblesup: Scribble-supervised convolutional networks for semantic segmentation. In CVPR, 2016.
  50. 50.Wei Liu, Andrew Rabinovich, and Alexander C. Berg. ParseNet: Looking Wider to See Better. arXiv:1506.04579, 2015.
  51. 51.Weide Liu, Chi Zhang, Guosheng Lin, and Fayao Liu. Crnet: Cross-reference networks for few-shot segmentation. In CVPR, 2020.
  52. 52.Timo Lüddecke and Alexander Ecker. Image segmentation using text and image prompts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7086–7096, 2022.
  53. 53.Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens Van Der Maaten. Exploring the limits of weakly supervised pretraining. In ECCV, 2018.
  54. 54.Kevis-Kokitsi Maninis, Sergi Caelles, Jordi Pont-Tuset, and Luc Van Gool. Deep extreme cut: From extreme points to object segmentation. In CVPR, 2018.
  55. 55.Luke Melas-Kyriazi, Christian Rupprecht, Iro Laina, and Andrea Vedaldi. Deep spectral methods: A surprisingly strong baseline for unsupervised semantic segmentation and localization. arXiv:2205.07839, 2022.
  56. 56.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. arXiv:1301.3781, 2013.
  57. 57.George A Miller. Wordnet: a lexical database for english. Communications of the ACM, 1995.
  58. 58.Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille. The role of context for object detection and semantic segmentation in the wild. In CVPR, 2014.
  59. 59.Lopamudra Mukherjee, Vikas Singh, and Charles R Dyer. Half-integrality based algorithms for cosegmentation of images. In CVPR, 2009.
  60. 60.Muzammal Naseer, Kanchana Ranasinghe, Salman Khan, Munawar Hayat, Fahad Khan, and Ming-Hsuan Yang. Intriguing properties of vision transformers. In NeurIPS, 2021.
  61. 61.Yassine Ouali, Céline Hudelot, and Myriam Tami. Autoregressive unsupervised image segmentation. In ECCV, 2020.
  62. 62.Fabio Panella, Victor Melatti, and Jan Boehm. Firenet dataset. http://www.firenet.xyz. Accessed: 2022-05-17.
  63. 63.Dim P Papadopoulos, Alasdair DF Clarke, Frank Keller, and Vittorio Ferrari. Training object class detectors from eye tracking data. In ECCV, 2014.
  64. 64.Dim P Papadopoulos, Jasper RR Uijlings, Frank Keller, and Vittorio Ferrari. Extreme clicking for efficient object annotation. In ICCV, 2017.
  65. 65.Giuseppe Pastore, Fabio Cermelli, Yongqin Xian, Massimiliano Mancini, Zeynep Akata, and Barbara Caputo. A closer look at self-training for zero-label semantic segmentation. In CVPR, 2021.
  66. 66.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, 2019.
  67. 67.Deepak Pathak, Philipp Krahenbuhl, and Trevor Darrell. Constrained convolutional neural networks for weakly supervised segmentation. In ICCV, 2015.
  68. 68.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In EMNLP, 2014.
  69. 69.Pedro O Pinheiro and Ronan Collobert. From image-level to pixel-level labeling with convolutional networks. In CVPR, 2015.
  70. 70.Rui Qian, Yunchao Wei, Honghui Shi, Jiachen Li, Jiaying Liu, and Thomas Huang. Weakly supervised scene parsing with point-based distance metric learning. In AAAI, 2019.
  71. 71.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In ICML, 2021.
  72. 72.Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu. Denseclip: Language-guided dense prediction with context-aware prompting. arXiv:2112.01518, 2021.
  73. 73.Carsten Rother, Tom Minka, Andrew Blake, and Vladimir Kolmogorov. Cosegmentation of image pairs by histogram matching-incorporating a global constraint into mrfs. In CVPR, 2006.
  74. 74.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. IJCV, 2015.
  75. 75.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Theo Coombes, Cade Gordon, Aarush Katta, Robert Kaczmarczyk, and Jenia Jitsev. LAION-5B: laion-5b: A new era of open large-scale multi-modal datasets. https://laion.ai/laion-5b-a-new-era-of-open-large-scale-multi-modal-datasets/, 2022.
  76. 76.Tong Shen, Guosheng Lin, Chunhua Shen, and Ian Reid. Bootstrapping the performance of webly supervised semantic segmentation. In CVPR, 2018.
  77. 77.Gyungin Shin, Samuel Albanie, and Weidi Xie. Unsupervised salient object detection with spectral cluster voting. In CVPRW, 2022.
  78. 78.Gyungin Shin, Weidi Xie, and Samuel Albanie. All you need are a few pixels: semantic segmentation with pixelpick. In ICCVW, 2021.
  79. 79.Oriane Siméoni, Gilles Puy, Huy V Vo, Simon Roburin, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Renaud Marlet, and Jean Ponce. Localizing objects with self-supervised transformers and no labels. arXiv:2109.14279, 2021.
  80. 80.Chunfeng Song, Yan Huang, Wanli Ouyang, and Liang Wang. Box-driven class-wise region masking and filling rate guided loss for weakly supervised semantic segmentation. In CVPR, 2019.
  81. 81.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, Vijay Vasudevan, Wei Han, Jiquan Ngiam, Hang Zhao, Aleksei Timofeev, Scott Ettinger, Maxim Krivokon, Amy Gao, Aditya Joshi, Yu Zhang, Jonathon Shlens, Zhifeng Chen, and Dragomir Anguelov. Scalability in perception for autonomous driving: Waymo open dataset. In CVPR, 2020.
  82. 82.Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, and Luc Van Gool. Unsupervised semantic segmentation by contrasting object mask proposals. In ICCV, 2021.
  83. 83.Sara Vicente, Vladimir Kolmogorov, and Carsten Rother. Cosegmentation revisited: Models and optimization. In ECCV, 2010.
  84. 84.Sara Vicente, Carsten Rother, and Vladimir Kolmogorov. Object cosegmentation. In CVPR, 2011.
  85. 85.Antonin Vobecky, David Hurych, Oriane Siméoni, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, and Josef Sivic. Drive&segment: Unsupervised semantic segmentation of urban scenes via cross-modal distillation. arXiv:2203.11160, 2022.
  86. 86.Yangtao Wang, Xi Shen, Shell Hu, Yuan Yuan, James Crowley, and Dominique Vaufreydaz. Self-supervised transformers for unsupervised object discovery using normalized cut. arXiv:2202.11539, 2022.
  87. 87.Mark Weber, Jun Xie, Maxwell Collins, Yukun Zhu, Paul Voigtlaender, Hartwig Adam, Bradley Green, Andreas Geiger, Bastian Leibe, Daniel Cremers, Aljosa Osep, Laura Leal-Taixe, and Liang-Chieh Chen. Step: Segmenting and tracking every pixel. In NeurIPS Track on Datasets and Benchmarks, 2021.
  88. 88.Yunchao Wei, Huaxin Xiao, Honghui Shi, Zequn Jie, Jiashi Feng, and Thomas S. Huang. Revisiting dilated convolution: A simple approach for weakly- and semi-supervised semantic segmentation. In CVPR, 2018.
  89. 89.Yongqin Xian, Subhabrata Choudhury, Yang He, Bernt Schiele, and Zeynep Akata. Semantic projection network for zero-and few-label semantic segmentation. In CVPR, 2019.
  90. 90.Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. Groupvit: Semantic segmentation emerges from text supervision. arXiv:2202.11094, 2022.
  91. 91.Mengde Xu, Zheng Zhang, Fangyun Wei, Yutong Lin, Yue Cao, Han Hu, and Xiang Bai. A simple baseline for zero-shot semantic segmentation with pre-trained vision-language model. arXiv:2112.14757, 2021.
  92. 92.Chi Zhang, Guankai Li, Guosheng Lin, Qingyao Wu, and Rui Yao. Cyclesegnet: Object co-segmentation with cycle refinement and region correspondence. IEEE TIP, 2021.
  93. 93.Xiao Zhang and Michael Maire. Self-supervised visual representation learning from hierarchical grouping. In NeurIPS, 2020.
  94. 94.Hang Zhao, Xavier Puig, Bolei Zhou, Sanja Fidler, and Antonio Torralba. Open vocabulary scene parsing. In ICCV, 2017.
  95. 95.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In CVPR, 2016.
  96. 96.Chong Zhou, Chen Change Loy, and Bo Dai. Denseclip: Extract free dense labels from clip. arXiv:2112.01071, 2021.

Citation

MLA
Shin, G., et al. “ReCo: Retrieve and Co-segment for Zero-shot Transfer”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 33754–67, https://proceedings.neurips.cc/paper_files/paper/2022/file/daabe43c3e1d06980aa23880bfbe1f45-Paper-Conference.pdf.
APA
Shin, G., Xie, W., & Albanie, S. (2022). ReCo: Retrieve and Co-segment for Zero-shot Transfer. Advances in Neural Information Processing Systems, 35, 33754–33767. https://proceedings.neurips.cc/paper_files/paper/2022/file/daabe43c3e1d06980aa23880bfbe1f45-Paper-Conference.pdf
Chicago
Shin, G., W. Xie, and S. Albanie. 2022. “ReCo: Retrieve and Co-segment for Zero-shot Transfer”. Advances in Neural Information Processing Systems 35: 33754–67. https://proceedings.neurips.cc/paper_files/paper/2022/file/daabe43c3e1d06980aa23880bfbe1f45-Paper-Conference.pdf.
Harvard
Shin, G., Xie, W. and Albanie, S. (2022) “ReCo: Retrieve and Co-segment for Zero-shot Transfer”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 33754–33767. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/daabe43c3e1d06980aa23880bfbe1f45-Paper-Conference.pdf.
Vancouver
1. Shin G, Xie W, Albanie S (2022) ReCo: Retrieve and Co-segment for Zero-shot Transfer. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 33754–33767

BibTeX

@inproceedings{shin2022reco,
  title = {ReCo: Retrieve and Co-segment for Zero-shot Transfer},
  author = {Shin, Gyungin and Xie, Weidi and Albanie, Samuel},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {33754-33767},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/daabe43c3e1d06980aa23880bfbe1f45-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors