CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor

Shuyang SunRunjia LiPhilip TorrXiuye GuSiyang Li

article2024CVPR96 citations

Presents a training-free recurrent framework that leverages frozen CLIP models to iteratively refine mask proposals and filter non-existent text queries, setting new state-of-the-art benchmarks in zero-shot open-vocabulary semantic and referring segmentation without sacrificing vocabulary breadth.

Listen

Modern computer vision increasingly requires systems to segment arbitrary visual concepts described by natural language, known as open-vocabulary segmentation. However, standard methods require fine-tuning pre-trained vision-language models on task-specific mask annotations or massive image-text datasets. This fine-tuning process is labor-intensive and dramatically degrades the extensive vocabulary inherited from base models, restricting their recognition capabilities on diverse concepts such as specific brands, landmarks, and fine-grained categories.

The article demonstrates that high-quality visual segmentation can be achieved directly from frozen pre-trained vision-language models without any fine-tuning or extra training. The primary objective is to evaluate whether a recurrent architecture can iteratively align textual queries with visual image features to eliminate irrelevant concepts and generate precise segmentation masks.

To accomplish this, the authors designed a training-free framework called CLIP as RNN. The framework processes an input image and a list of unrestricted text queries through an iterative two-stage cycle using frozen model weights. In each recurrent step, a proposal generator creates candidate visual masks, and a classifier evaluates visual-textual alignment using visual prompts (such as background blur and red circles) to progressively filter out unmatched or nonexistent queries. The process repeats until the query set stabilizes, followed by standard boundary refinement. The methodology was evaluated across eight standard semantic and referring segmentation benchmarks.

Across zero-shot semantic segmentation benchmarks, the proposed method achieved major performance gains without fine-tuning. Compared to existing training-free techniques, it improved mean Intersection-over-Union by 28.8 points on Pascal VOC, 16.0 points on COCO Object, and 6.9 points on Pascal Context. Notably, it also outperformed competing approaches that had been fine-tuned on tens to hundreds of millions of images, surpassing the top fine-tuned baseline by 12.6 points on Pascal VOC and 4.6 points on COCO Object. On referring expression benchmarks, the method established new state-of-the-art zero-shot accuracy across RefCOCO, RefCOCO+, and RefCOCOg, and created a competitive zero-shot baseline for referring video segmentation.

These findings prove that costly data collection, annotation, and model retraining pipelines are not strictly necessary for advanced visual segmentation. By retaining the original weights of foundation models, organizations can preserve extensive open vocabularies while lowering computational overhead and engineering complexity. The approach operated efficiently on a single standard graphics processor, requiring minimal memory and execution time.

Organizations developing computer vision applications should consider adopting recurrent inference pipelines before investing in expensive fine-tuning workflows for open-vocabulary tasks. For production implementations requiring maximum edge precision, integrating optional post-processing segmenters can yield additional boundary accuracy.

The evidence presented provides high confidence in the framework's effectiveness across common semantic and referring segmentation datasets. However, the source notes that the model exhibits reduced sensitivity when segmenting broad contextual background regions (such as "stuff" categories like sky or grass) because foundation models encounter these concepts less frequently during initial training.

Cover for CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor

Abstract

Existing open-vocabulary image segmentation methods require a fine-tuning step on mask labels and/or image-text datasets. Mask labels are labor-intensive, which limits the number of categories in segmentation datasets. Consequently, the vocabulary capacity of pre-trained VLMs is severely reduced after fine-tuning. However, without fine-tuning, VLMs trained under weak image-text supervision tend to make suboptimal mask predictions. To alleviate these issues, we introduce a novel recurrent framework that progressively filters out irrelevant texts and enhances mask quality without training efforts. The recurrent unit is a two-stage segmenter built upon a frozen VLM. Thus, our model retains the VLM’s broad vocabulary space and equips it with segmentation ability. Experiments show that our method outperforms not only the training-free counterparts, but also those fine-tuned with millions of data samples, and sets the new state-of-the-art records for both zero-shot semantic and referring segmentation. Concretely, we improve the current record by 28.8, 16.0, and 6.9 mIoU on Pascal VOC, COCO Object, and Pascal Context.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. CLIP as Recurrent Neural Networks
  • 3.1. A Recap on Recurrent Neural Networks
  • 3.2. Overview
  • 3.3. The Two-stage Segmenter
  • 3.4. Post-Processing
  • 4. Experiments
  • 4.1. Zero-shot Semantic Segmentation
  • 4.2. Ablation Studies
  • 4.3. Referring Segmentation
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — CaR preserves CLIP’s vocabulary through recurrent inference

    model/method

    CLIP as RNN (CaR) is a training-free open-vocabulary segmentation framework that keeps the weights and architecture of a pre-trained vision-language model (VLM) frozen. Given an image xt∈R3×H×Wx_t \in \mathbb{R}^{3\times H\times W} and an initial list of arbitrary text queries h0h_0—including object classes, proper nouns, fictional characters, landmarks, brands, or referring expressions—CaR repeatedly applies one shared two-stage segmenter. At time step tt, the segmenter receives the image and the surviving query list ht−1h_{t-1}, produces one mask proposal per query, evaluates the visual-textual alignment of each proposal, and removes low-confidence queries. For a static image, xt=xx_t=x at every time step. The process stops when the query list is unchanged, ht=ht−1h_t=h_{t-1}, and the masks from the terminal step are returned. Because no mask labels, image-text pairs, or other data are used for fine-tuning, the broad vocabulary encoded by the original VLM is retained.

  2. Knowl 2 — Frozen CLIP mask proposal generator with explicit background queries

    model/method

    For each surviving text query, CaR’s mask proposal generator produces a dense candidate mask using gradient-based class activation maps from frozen CLIP features. The image and query are first scored by CLIP; the gradient of that query score is back-propagated through the CLIP image encoder to obtain a spatial heatmap. The implementation uses the CAM and class-affinity (CAA) components of CLIP-ES as the proposal generator, without training CLIP. In addition to the user’s queries, CaR supplies background queries describing categories absent from the user query list. These background categories are grouped into terrestrial, aquatic, atmospheric, and man-made concepts, and their gradients help suppress activations caused by irrelevant text queries. The proposal generator is represented by ff with frozen weights WfW_f and computes

    yt=f(xt,ht−1;Wf),y_t=f(x_t,h_{t-1};W_f),

    where ht−1h_{t-1} contains Nt−1N_{t-1} text queries and yt∈[0,1]Nt−1×H×Wy_t\in[0,1]^{N_{t-1}\times H\times W} contains one soft mask proposal for each query.

  3. Knowl 3 — Visual-prompt mask classifier filters irrelevant queries

    equation

    CaR assesses each proposal by converting it into a visually prompted image rather than classifying a bare foreground crop. Let vv be the visual-prompting function and let η∈[0,1]\eta\in[0,1] be the threshold used to binarize each soft proposal. The resulting Nt−1N_{t-1} prompted images are

    xt′=v(xt,yt).x'_t=v(x_t,y_t).

    The prompts can draw a red circle or contour around the proposal, blur or gray the background, or mask the background while retaining some context. A frozen CLIP classifier gg with weights WgW_g compares every prompted image with every query and produces a normalized similarity matrix

    Pt=g(xt′,ht−1;Wg)∈RNt−1×Nt−1,P_t=g(x'_t,h_{t-1};W_g)\in\mathbb{R}^{N_{t-1}\times N_{t-1}},

    where PtiiP_t^{ii} is the softmax-normalized score for the ii-th prompted mask paired with its own ii-th query. The query is retained only when its diagonal score exceeds a manually selected threshold θ\theta:

    h_t^i=\begin{cases}h_{t-1}^i,&P_t^{ii}\geq\theta,\\\mathrm{NULL},&P_t^{ii}<\theta.\end{cases}$$ Thus, queries referring to concepts unsupported by the image are progressively removed, causing later proposal-generation steps to operate on a cleaner query set.
  4. Knowl 4 — Inference procedure for CLIP as RNN

    algorithm

    CaR inference takes an image, an initial query list, frozen CLIP models, a CAM-based proposal generator, a mask-binarization threshold η\eta, and a query-retention threshold θ\theta. It returns the proposals from the final recurrent step, followed by post-processing.

    Input: image x, initial text-query list h_0, frozen CLIP, CAM proposal generator, thresholds eta and theta
    Output: final segmentation masks
    h = h_0
    while h is not empty:
        Compute CLIP image-text logits between x and every query in h
        Normalize the logits over the query dimension
        Generate one CAM/CAA mask proposal for each query, also using background queries
        Binarize the proposals at threshold eta
        Create one visually prompted image for each binarized proposal
        Compute CLIP logits between every prompted image and every query in h
        Normalize the logits over the query dimension
        For each query, read the diagonal score of its own prompted-image/query pair
        Keep the query if its diagonal score is at least theta
        If every query was kept, stop the recurrent loop
        Otherwise, replace h with the retained query list
    Apply dense-CRF post-processing to the proposals from the last recurrent step
    Return the resulting masks
  5. Knowl 5 — Dense-CRF and optional SAM boundary refinement

    model/method

    After recurrence terminates at time TT, CaR refines the final proposals yTy_T with a dense conditional random field (CRF). The CRF unary potentials are derived from the final proposal responses, its hyperparameters use the defaults of the dense-CRF implementation adopted by the paper, and an argmax over the text-query dimension assigns one query label to each pixel. CaR also supports an optional SAM-based refinement that does not provide prompts to SAM: SAM is run in automask mode to generate candidate masks, which are matched to CRF masks using the Intersection over Minimum-mask metric. For binary pixel sets AA and BB, this metric is

    IoM⁡(A,B)=∣A∩B∣min⁡(∣A∣,∣B∣).\operatorname{IoM}(A,B)=\frac{|A\cap B|}{\min(|A|,|B|)}.

    SAM masks with IoM⁡(A,B)>ϕiom\operatorname{IoM}(A,B)>\phi_{\mathrm{iom}} are merged for the same CRF mask. The merged mask replaces the original CRF mask only if its IoU with the original mask exceeds ϕiou\phi_{\mathrm{iou}}; otherwise the CRF mask is retained. SAM is therefore used only as an optional post-processor, not as a trained segmentation component.

  6. Knowl 6 — Experimental configuration and inference cost

    experimental setup

    The semantic-segmentation evaluation uses validation splits of Pascal VOC, COCO Object, and Pascal Context, with additional evaluation on ADE-150, ADE-847, and Pascal Context 459. COCO Object contains 80 object categories plus one merged background class. Pascal VOC is evaluated both without background (VOC-20) and with background (VOC-21); Pascal Context is similarly evaluated as PC-59 and PC-60. The primary metric is mean intersection-over-union (mIoU).

    CaR uses CLIP ViT-B/16 for the mask proposal generator and the larger ViT-L/14 for the mask classifier. Unless SAM is explicitly enabled, results use dense-CRF only. The paper uses (η,θ,λ)=(0.4,0.6,0.4)(\eta,\theta,\lambda)=(0.4,0.6,0.4) for Pascal VOC, (0.5,0.3,0.5)(0.5,0.3,0.5) for COCO Object, and (0.6,0.2,0.4)(0.6,0.2,0.4) for Pascal Context, where η\eta binarizes proposals, θ\theta filters queries, and λ\lambda is the CLIP-ES proposal parameter. Optional SAM refinement uses ϕiom=ϕiou=0.7\phi_{\mathrm{iom}}=\phi_{\mathrm{iou}}=0.7. CLIP inference uses half precision. On one NVIDIA V100 GPU, CaR without SAM takes about 950 ms for a 500×500500\times500 image when using a CPU dense CRF; a GPU CRF is approximately five times faster, and Pascal VOC inference uses 3.6 GB of GPU memory.

  7. Knowl 7 — Training-free semantic segmentation substantially exceeds prior methods

    data/table

    CaR achieves the following mIoU results without fine-tuning CLIP, auxiliary segmentation training, or additional image-text data. The comparison margins are against the best reported competing method in the corresponding category; a dash indicates that the paper does not report a margin for that dataset. The semantic benchmarks differ in whether background is included, as indicated by their names.

    Could not parse LaTeX table

    These results show that recurrent inference with frozen CLIP outperforms both earlier training-free methods and several methods fine-tuned using millions of additional examples. Adding optional SAM post-processing further improves CaR by 2.6 mIoU on Pascal VOC, 1.1 on COCO Object, and 0.6 on Pascal Context.

  8. Knowl 8 — Ablations show recurrence and the asymmetric CLIP backbone choice are load-bearing

    data/table

    Ablation experiments on Pascal VOC show that recurrence, the larger classifier backbone, contextual visual prompts, and diverse background queries each affect CaR’s performance. Without recurrence, the CLIP-ES-based system obtains only 15.2 mIoU; adding the recurrent filtering process raises this to 67.6 mIoU. Replacing the CAM proposal method with plain gradCAM while retaining recurrence gives 41.1 mIoU, showing that the recurrent design and the proposal quality both matter.

    For the mask proposal generator ff and classifier gg, the reported mIoU values are:

    Could not parse LaTeX table

    The best single visual prompt on Pascal VOC is the red-circle prompt at 66.9 mIoU, whereas masking the background alone gives 61.8; combining a red circle with background blur gives the best reported prompt result, 67.6 mIoU. Background-query diversity also matters: using no background queries gives 64.3 mIoU, while combining terrestrial, aquatic, atmospheric, and man-made background-query groups gives 67.6 mIoU.

  9. Knowl 9 — CaR improves zero-shot referring image segmentation

    data/table

    For referring image segmentation, CaR takes a natural-language expression identifying one region and performs inference without training or fine-tuning on referring annotations. The reported results use dense-CRF post-processing only; SAM is not used. The metric is mIoU, and the comparison is against the zero-shot Global-Local CLIP (GL CLIP) method.

    Could not parse LaTeX table

    CaR is higher than GL CLIP on every reported split. The largest gains are 10.42 mIoU on RefCOCO testA and 10.72 mIoU on RefCOCO+ testA. CaR also provides the reported first zero-shot result on GRES, achieving 16.8 mIoU.

  10. Knowl 10 — Zero-shot video referring segmentation baseline

    empirical result

    The paper extends the same inference-only CaR framework to referring video segmentation on Ref-DAVIS 2017, without fine-tuning or annotation-specific training. Using region similarity J\mathcal{J}, contour accuracy F\mathcal{F}, and their mean J&F\mathcal{J}\&\mathcal{F}, CaR obtains J&F=30.34\mathcal{J}\&\mathcal{F}=30.34, J=28.15\mathcal{J}=28.15, and F=32.53\mathcal{F}=32.53. These results establish a zero-shot baseline for language-guided video segmentation under the paper’s frozen-VLM setting.

Coverage note — No substantial contributed material was omitted; lower-level CAM/CAA implementation details and supplementary prompt lists were condensed because they are implementation refinements rather than separate load-bearing contributions.

References

  1. 1.Nikita Araslanov and Stefan Roth. Single-stage semantic segmentation from image labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4253–4262, 2020.
  2. 2.Donghyeon Baek, Youngmin Oh, and Bumsub Ham. Exploiting a joint embedding space for generalized zero-shot semantic segmentation. In ICCV, 2021.
  3. 3.Maxime Bucher, Tuan-Hung Vu, Matthieu Cord, and Patrick Perez. Zero-shot semantic segmentation. Advances in Neural Information Processing Systems, 32, 2019.
  4. 4.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2018.
  5. 5.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1209–1218, 2018.
  6. 6.Kaixin Cai, Pengzhen Ren, Yi Zhu, Hang Xu, Jianzhuang Liu, Changlin Li, Guangrun Wang, and Xiaodan Liang. Mixreorg: Cross-modal mixed patch reorganization is a good mask learner for open-world semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1196–1205, 2023.
  7. 7.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In CVPR, 2018.
  8. 8.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020.
  9. 9.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021.
  10. 10.Junbum Cha, Jonghwan Mun, and Byungseok Roh. Learning to generate text-grounded mask for open-world semantic segmentation from only image-text pairs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11165–11174, 2023.
  11. 11.Jun Chen, Deyao Zhu, Guocheng Qian, Bernard Ghanem, Zhicheng Yan, Chenchen Zhu, Fanyi Xiao, Mohamed Elhoseiny, and Sean Chang Culatana. Exploring open-vocabulary semantic segmentation without human labels. arXiv preprint arXiv:2306.00450, 2023.
  12. 12.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. In ICLR, 2015.
  13. 13.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018.
  14. 14.Peijie Chen, Qi Li, Saad Biaz, Trung Bui, and Anh Nguyen. gscorecam: What objects is clip looking at? In Proceedings of the Asian Conference on Computer Vision, pages 1959–1975, 2022.
  15. 15.Bowen Cheng, Alexander G Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. In NeurIPS, 2021.
  16. 16.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. CVPR, 2022.
  17. 17.Jian Ding, Nan Xue, Gui-Song Xia, and Dengxin Dai. Decoupling zero-shot semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11583–11592, 2022.
  18. 18.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. IJCV, 88:303–338, 2010.
  19. 19.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88:303–338, 2010.
  20. 20.Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin. Scaling open-vocabulary image segmentation with image-level labels. In European Conference on Computer Vision, pages 540–557. Springer, 2022.
  21. 21.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 580–587, 2014.
  22. 22.Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open-vocabulary object detection via vision and language knowledge distillation. arXiv preprint arXiv:2104.13921, 2021.
  23. 23.Wenbin He, Suphanut Jamonnak, Liang Gou, and Liu Ren. Clip-s4: Language-guided self-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11207–11216, 2023.
  24. 24.Ping Hu, Stan Sclaroff, and Kate Saenko. Uncertainty-aware learning for zero-shot semantic segmentation. Advances in Neural Information Processing Systems, 33:21713–21724, 2020.
  25. 25.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning, pages 4904–4916. PMLR, 2021.
  26. 26.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. Mdetr-modulated detection for end-to-end multi-modal understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1780–1790, 2021.
  27. 27.Laurynas Karazija, Iro Laina, Andrea Vedaldi, and Christian Rupprecht. Diffusion models for zero-shot open-vocabulary segmentation. arXiv preprint arXiv:2306.09316, 2023.
  28. 28.Lei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu, Yu-Wing Tai, Chi-Keung Tang, and Fisher Yu. Segment anything in high quality. arXiv preprint arXiv:2306.01567, 2023.
  29. 29.Anna Khoreva, Anna Rohrbach, and Bernt Schiele. Video object segmentation with language referring expressions. In Computer Vision–ACCV 2018: 14th Asian Conference on Computer Vision, Perth, Australia, December 2–6, 2018, Revised Selected Papers, Part IV 14, pages 123–141. Springer, 2019.
  30. 30.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023.
  31. 31.Philipp Krahenb¨uhl and Vladlen Koltun. Efficient inference in fully connected crfs with gaussian edge potentials. Advances in neural information processing systems, 24, 2011.
  32. 32.Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic segmentation. In International Conference on Learning Representations, 2022.
  33. 33.Peike Li, Yunchao Wei, and Yi Yang. Consistent structural relation learning for zero-shot segmentation. NeurIPS, 2020.
  34. 34.Yanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer, and Kaiming He. Scaling language-image pre-training via masking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23390–23400, 2023.
  35. 35.Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu. Open-vocabulary semantic segmentation with mask-adapted clip. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7061–7070, 2023.
  36. 36.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740–755. Springer, 2014.
  37. 37.Yuqi Lin, Minghao Chen, Wenxiao Wang, Boxi Wu, Ke Li, Binbin Lin, Haifeng Liu, and Xiaofei He. Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15305–15314, 2023.
  38. 38.Chang Liu, Henghui Ding, and Xudong Jiang. GRES: Generalized referring expression segmentation. In CVPR, 2023.
  39. 39.Quande Liu, Youpeng Wen, Jianhua Han, Chunjing Xu, Hang Xu, and Xiaodan Liang. Open-world semantic segmentation via contrasting and clustering vision-language embedding. In European Conference on Computer Vision, pages 275–292. Springer, 2022.
  40. 40.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023.
  41. 41.Timo Luddecke and Alexander Ecker. Image segmentation using text and image prompts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7086–7096, 2022.
  42. 42.Huaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He, and Tianrui Li. Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation. In International Conference on Machine Learning, pages 23033–23044. PMLR, 2023.
  43. 43.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 11–20, 2016.
  44. 44.M Minderer, A Gritsenko, A Stone, M Neumann, D Weissenborn, A Dosovitskiy, A Mahendran, A Arnab, M Dehghani, Z Shen, et al. Simple open-vocabulary object detection with vision transformers. arxiv 2022. arXiv preprint arXiv:2205.06230, 2022.
  45. 45.Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille. The role of context for object detection and semantic segmentation in the wild. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014.
  46. 46.Jishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed, Rui Wang, Ashish Shah, Philip HS Torr, and Ser-Nam Lim. Open vocabulary semantic segmentation with patch aligned contrastive learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19413–19423, 2023.
  47. 47.Varun K Nagaraja, Vlad I Morariu, and Larry S Davis. Modeling context between objects for referring expression understanding. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14, pages 792–807. Springer, 2016.
  48. 48.Jianmo Ni, Gustavo Hernández Abrego, Noah Constant, Ji Ma, Keith B Hall, Daniel Cer, and Yinfei Yang. Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models. arXiv preprint arXiv:2108.08877, 2021.
  49. 49.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  50. 50.Kanchana Ranasinghe, Brandon McKinzie, Sachin Ravi, Yinfei Yang, Alexander Toshev, and Jonathon Shlens. Perceptual grouping in vision-language models. arXiv preprint arXiv:2210.09996, 2022.
  51. 51.Pengzhen Ren, Changlin Li, Hang Xu, Yi Zhu, Guangrun Wang, Jianzhuang Liu, Xiaojun Chang, and Xiaodan Liang. Viewco: Discovering text-supervised segmentation masks via multi-view semantic consistency. arXiv preprint arXiv:2302.10307, 2023.
  52. 52.Lixiang Ru, Yibing Zhan, Baosheng Yu, and Bo Du. Learning affinity from attention: End-to-end weakly-supervised semantic segmentation with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16846–16855, 2022.
  53. 53.Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, pages 618–626, 2017.
  54. 54.Gyungin Shin, Weidi Xie, and Samuel Albanie. Reco: Retrieve and co-segment for zero-shot transfer. Advances in Neural Information Processing Systems, 35:33754–33767, 2022.
  55. 55.Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi. What does clip know about a red circle? visual prompt engineering for vlms. arXiv preprint arXiv:2304.06712, 2023.
  56. 56.Robin Strudel, Ivan Laptev, and Cordelia Schmid. Weakly-supervised segmentation of referring expressions. arXiv preprint arXiv:2205.04725, 2022.
  57. 57.Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023.
  58. 58.Shuyang Sun, Weijun Wang, Qihang Yu, Andrew Howard, Philip Torr, and Liang-Chieh Chen. Remax: Relaxing for better training on efficient panoptic segmentation. arXiv preprint arXiv:2306.17319, 2023.
  59. 59.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  60. 60.Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Max-deeplab: End-to-end panoptic segmentation with mask transformers. In CVPR, 2021.
  61. 61.Haoxiang Wang, Pavan Kumar Anasosalu Vasu, Fartash Faghri, Raviteja Vemulapalli, Mehrdad Farajtabar, Sachin Mehta, Mohammad Rastegari, Oncel Tuzel, and Hadi Pouransari. Sam-clip: Merging vision foundation models towards semantic and spatial understanding. arXiv preprint arXiv:2310.15308, 2023.
  62. 62.Xinlong Wang, Zhiding Yu, Shalini De Mello, Jan Kautz, Anima Anandkumar, Chunhua Shen, and Jose M Alvarez. Freesolo: Learning to segment objects without annotations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14176–14186, 2022.
  63. 63.Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv:1609.08144, 2016.
  64. 64.Yongqin Xian, Subhabrata Choudhury, Yang He, Bernt Schiele, and Zeynep Akata. Semantic projection network for zero-and few-label semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8256–8265, 2019.
  65. 65.Jinheng Xie, Xianxu Hou, Kai Ye, and Linlin Shen. Clims: Cross language image matching for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4483–4492, 2022.
  66. 66.Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. Groupvit: Semantic segmentation emerges from text supervision. In CVPR, 2022.
  67. 67.Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng, Yi Wang, Yu Qiao, and Weidi Xie. Learning open-vocabulary semantic segmentation models from natural language supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2935–2944, 2023.
  68. 68.Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xiaolong Wang, and Shalini De Mello. Open-vocabulary panoptic segmentation with text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2955–2966, 2023.
  69. 69.Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaid, and Dan Xu. Multi-class token transformer for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4310–4319, 2022.
  70. 70.Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, and Philip HS Torr. Lavt: Language-aware vision transformer for referring image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18155–18165, 2022.
  71. 71.Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917, 2022.
  72. 72.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling context in referring expressions. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14, pages 69–85. Springer, 2016.
  73. 73.Qihang Yu, Huiyu Wang, Siyuan Qiao, Maxwell Collins, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. k-means Mask Transformer. In ECCV, 2022.
  74. 74.Qihang Yu, Ju He, Xueqing Deng, Xiaohui Shen, and Liang-Chieh Chen. Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip. arXiv preprint arXiv:2308.02487, 2023.
  75. 75.Seonghoon Yu, Paul Hongsuck Seo, and Jeany Son. Zero-shot referring image segmentation with global-local context features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19456–19465, 2023.
  76. 76.Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer. Lit: Zero-shot transfer with locked-image text tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18123–18133, 2022.
  77. 77.Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. arXiv preprint arXiv:2303.15343, 2023.
  78. 78.Haotian Zhang, Pengchuan Zhang, Xiaowei Hu, Yen-Chun Chen, Liunian Li, Xiyang Dai, Lijuan Wang, Lu Yuan, Jenq-Neng Hwang, and Jianfeng Gao. Glipv2: Unifying localization and vision-language understanding. Advances in Neural Information Processing Systems, 35:36067–36080, 2022.
  79. 79.Hao Zhang, Feng Li, Xueyan Zou, Shilong Liu, Chunyuan Li, Jianwei Yang, and Lei Zhang. A simple framework for open-vocabulary segmentation and detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1020–1031, 2023.
  80. 80.Shuai Zheng, Sadeep Jayasumana, Bernardino Romera-Paredes, Vibhav Vineet, Zhizhong Su, Dalong Du, Chang Huang, and Philip HS Torr. Conditional random fields as recurrent neural networks. In ICCV, 2015.
  81. 81.Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ade20k dataset. International Journal of Computer Vision, 2019.
  82. 82.Chong Zhou, Chen Change Loy, and Bo Dai. Extract free dense labels from clip. In European Conference on Computer Vision, pages 696–712. Springer, 2022.
  83. 83.Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krahenb¨uhl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. In European Conference on Computer Vision, pages 350–368. Springer, 2022.

Citation

MLA
Sun, S., et al. “CLIP as RNN: Segment Countless Visual Concepts Without Training Endeavor”. arXiv, 2023, http://arxiv.org/abs/2312.07661v3.
APA
Sun, S., Li, R., Torr, P., Gu, X., & Li, S. (2023). CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor. arXiv. http://arxiv.org/abs/2312.07661v3
Chicago
Sun, S., R. Li, P. Torr, X. Gu, and S. Li. 2023. “CLIP as RNN: Segment Countless Visual Concepts Without Training Endeavor”. arXiv. http://arxiv.org/abs/2312.07661v3.
Harvard
Sun, S. et al. (2023) “CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.07661v3.
Vancouver
1. Sun S, Li R, Torr P, Gu X, Li S (2023) CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor. arXiv

BibTeX

@article{sun2023clip,
  title = {CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor},
  author = {Sun, Shuyang and Li, Runjia and Torr, Philip and Gu, Xiuye and Li, Siyang},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.07661v3},
  eprint = {2312.07661}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE