Alpha-CLIP: A CLIP Model Focusing on Wherever you Want

Zeyi SunYe FangTong WuPan ZhangYuhang ZangShu KongYuanjun XiongDahua LinJiaqi Wang

article2024CVPR218 citations

Introduces Alpha-CLIP, an enhanced CLIP model with an auxiliary alpha channel that enables fine-grained, region-specific focus while preserving contextual awareness across open-world recognition, multimodal language models, and 2D/3D generation tasks.

Listen

Modern artificial intelligence systems increasingly rely on vision-language models to interpret visual content and generate text, images, or three-dimensional assets. While conventional vision backbones excel at capturing entire scenes, they struggle to focus on specific regions of interest indicated by points, strokes, or masks. Existing workarounds, such as cropping objects or drawing visible outlines, either destroy vital surrounding context or alter the underlying image, leading to recognition errors and visual distortions. The article addresses this operational challenge by demonstrating a method to enable precise, region-specific focus in vision models without sacrificing global image context.

To achieve this, the article introduces Alpha-CLIP, an enhanced version of the standard vision backbone that incorporates an auxiliary alpha channel input alongside standard color channels. The researchers developed an automated pipeline that constructed millions of region-text pairs using segmentation and captioning tools, eliminating the need for costly manual annotations. They then trained the modified image encoder on this data using a sampling strategy that preserved full-image understanding while teaching the model to focus on designated regions.

The findings show that Alpha-CLIP substantially outperforms standard baselines across diverse applications. In zero-shot image classification, providing a foreground mask boosted top-one accuracy by approximately 4 percentage points, reaching 77.41% compared to 73.48% for the standard baseline. In referring expression tasks, it outperformed existing approaches by an average of 3 to 6.8 percentage points. When serving as a visual engine for object detection, the system achieved higher detection accuracy while using fewer than half the training samples of previous pipelines. Furthermore, integrating Alpha-CLIP into multimodal language frameworks significantly reduced factual hallucinations, while its application in 2D and 3D generative workflows produced cleaner, better-aligned shapes and sped up 3D optimization by two times.

These results indicate that organizations deploying vision and multimodal systems can achieve higher precision, reduced hallucination risks, and improved generative quality without re-engineering their core pipelines. Because Alpha-CLIP functions as a drop-in replacement that maintains output consistency with standard architectures, implementation costs and transition risks remain minimal. Practitioners in visual editing, automated inspection, and conversational visual interfaces should pilot Alpha-CLIP as an alternative vision backbone to evaluate gains in downstream task accuracy.

The primary operational limitation is that the model's enhanced capabilities depend on obtaining reasonable region proposals or masks from users or upstream segmentation tools, although it retains baseline performance when no region is specified. Overall, the consistent improvements demonstrated across standard benchmarks provide high confidence in Alpha-CLIP as a versatile, plug-and-play enhancement for fine-grained computer vision tasks.

arXiv: 2312.03818
Cover for Alpha-CLIP: A CLIP Model Focusing on Wherever you Want

Abstract

Contrastive Language-Image Pre-training (CLIP) plays an essential role in extracting valuable content information from images across diverse tasks. It aligns textual and visual modalities to comprehend the entire image, including all the details, even those irrelevant to specific tasks. However, for a finer understanding and controlled editing of images, it becomes crucial to focus on specific regions of interest, which can be indicated as points, masks, or boxes by humans or perception models. To fulfill the requirements, we introduce Alpha-CLIP, an enhanced version of CLIP with an auxiliary alpha channel to suggest attentive regions and fine-tuned with constructed millions of RGBA region-text pairs. Alpha-CLIP not only preserves the visual recognition ability of CLIP but also enables precise control over the emphasis of image contents. It demonstrates effectiveness in various tasks, including but not limited to open-world recognition, multimodal large language models, and conditional 2D / 3D generation. It has a strong potential to serve as a versatile tool for image-related tasks. Our project is with codes and models available is linked to https://aleafy.github.io/alpha-clip/.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. RGBA Region-Text Pair Generation
  • 3.2. Alpha-CLIP
  • 4. Experiments
  • 4.1. Alpha-CLIP in Image Recognition
  • 4.2. Alpha-CLIP in MLLM
  • 4.3. Alpha-CLIP in image variation.
  • 4.4. Alpha-CLIP in 3D Object Generation.
  • 5. Conclusion
  • 6. Acknowledgment
  • References

Knowls

  1. Knowl 1 — Alpha-channel region focusing without changing image content

    model/method

    Alpha-CLIP extends a pretrained CLIP image encoder with an auxiliary alpha channel that specifies the region of interest while leaving the RGB image unchanged. The alpha map has the same spatial resolution as the image and uses values in [0,1][0,1], where 11 denotes foreground or emphasized content and 00 denotes background. Unlike cropping, pixel masking, or drawing a visual marker onto the image, this design preserves contextual information and supports pixel-level region selection. The resulting image embeddings remain compatible with CLIP-based downstream systems, allowing Alpha-CLIP to replace the original CLIP image encoder in recognition, multimodal language, diffusion, and 3D-generation pipelines.

  2. Knowl 2 — Alpha-CLIP image-encoder architecture

    model/method

    Alpha-CLIP modifies the first layer of CLIP’s ViT image encoder by adding an alpha convolution in parallel with the original RGB convolution. The RGB image is processed by the inherited RGB convolution, the alpha map is processed by the new alpha convolution, and the two outputs are combined before entering the unchanged transformer stack. The alpha-convolution kernel is initialized to zero, so the initialized model ignores the alpha channel and starts from the behavior of the pretrained CLIP image encoder. The CLIP text encoder is retained, making the model a minimally modified, plug-in-compatible CLIP variant. The architecture diagram on page 4 shows the parallel RGB and alpha branches feeding the same attention blocks.

  3. Knowl 3 — Large-scale RGBA region-text data engine

    algorithm

    The paper constructs millions of RGBA region-text pairs through two complementary pipelines.

    Input: Grounding images with box-text annotations, ImageNet images, SAM, CLIP, and BLIP-2
    Output: RGBA images paired with region descriptions
    For each grounding image and each annotated box-text pair:
        Apply SAM to the box to obtain a foreground pseudo-mask.
        Store the original RGB image, the pseudo-mask as its alpha channel, and the region text.
    For each ImageNet image:
        Apply SAM to generate several candidate masks.
        Crop, center, and enlarge the foreground object for each candidate mask.
        Use CLIP to score each candidate against the image's ImageNet class label.
        Retain the highest-scoring masks for each class.
        Place each retained foreground object on a pure white background.
        Use BLIP-2 to generate an image-specific caption.
        Combine the fine-grained ImageNet class label with the BLIP-2 caption.
        Store the resulting RGBA image and combined text description.
    Return the grounding-derived and classification-derived RGBA region-text pairs.

    The grounding branch starts from GRIT-20m box-region annotations and uses SAM to convert boxes into mask-level annotations. The classification branch uses ImageNet masks, CLIP-based mask ranking, and BLIP-2 captions so that descriptions contain more information than an ImageNet class name alone. For the experiments beyond zero-shot ImageNet recognition, the authors use 460,000 ImageNet-derived RGBA pairs in addition to the grounding data.

  4. Knowl 4 — Mixed-data fine-tuning that preserves global CLIP recognition

    model/method

    Alpha-CLIP is initialized from CLIP, with the text encoder kept fixed and the image encoder trained using the generated RGBA region-text pairs. The newly added first-layer alpha convolution receives a higher learning rate than the subsequent transformer blocks, whose learning rate is reduced to limit disruption of CLIP’s pretrained representation. To retain whole-image recognition, a sample ratio rs=0.1r_s=0.1 is used: approximately 10% of training examples are replaced by original image-text pairs, with the alpha map set uniformly to 11. Thus, the model learns region-conditioned alignment while periodically training on ordinary full-image inputs.

  5. Knowl 5 — Zero-shot recognition improves with foreground alpha maps

    data/table

    On the ImageNet-S validation set, which contains 919 ImageNet classes with semantic segmentation annotations, the authors evaluate mean per-class zero-shot accuracy when the foreground mask is supplied as the alpha channel. Alpha-CLIP outperforms the original CLIP and the compared region-focusing baselines for both ViT-B/16 and ViT-L/14. The results reported on page 5 are:

    Could not parse LaTeX table

    The alpha-map ablation on the same benchmark shows that Alpha-CLIP remains comparable to CLIP without a foreground prior and improves as the region prior becomes more precise:

    Could not parse LaTeX table

    These results demonstrate that the model retains ordinary CLIP recognition when the alpha map is uninformative and gains region-focused recognition from box- or mask-level input.

  6. Knowl 6 — Zero-shot referring-expression comprehension with original context preserved

    empirical result

    For zero-shot referring-expression comprehension, Alpha-CLIP replaces CLIP in a proposal-based pipeline. A pretrained detector supplies object proposals, SAM converts each proposal into a mask, and the original image plus the proposal mask is passed to Alpha-CLIP instead of cropping the proposal out of the image. This preserves global context, which the authors report is beneficial compared with cropping. The reported top-1 accuracies are:

    Could not parse LaTeX table

    Alpha-CLIP exceeds ReCLIP by an average of 6.8 percentage points and Red Circle by an average of 3.0 percentage points across the RefCOCO, RefCOCO+, and RefCOCOg benchmarks, according to the paper’s aggregate comparison.

  7. Knowl 7 — MaskImageNet improves open-vocabulary detection with less data

    empirical result

    The authors use Alpha-CLIP in Detic’s image-level pseudo-labeling pipeline for open-vocabulary detection on OV-LVIS. They convert 460,000 top-ranked ImageNet images into MaskImageNet by generating pseudo bounding boxes and foreground masks, then replace the 1.2 million ImageNet images used by the original Detic pipeline. The reported detection results are:

    Could not parse LaTeX table

    MaskImageNet alone improves novel-class detection over the Detic baseline, while replacing the original CLIP with Alpha-CLIP provides a further increase. The Alpha-CLIP configuration reaches 28.6 novel-class mAP and 32.9 overall mAP while using 460,000 images rather than Detic’s 1.2 million-image ImageNet collection.

  8. Knowl 8 — Region-focused multimodal language understanding

    empirical result

    Replacing the CLIP image encoder in BLIP-2 and LLaVA-1.5 with Alpha-CLIP lets a multimodal large language model use an alpha map as a visual prompt. Users can identify a target object with a mask or stroke, after which captioning and visual question answering are conditioned more strongly on that region. In qualitative examples on page 7, ordinary CLIP mixes properties from multiple objects and produces incorrect captions, whereas Alpha-CLIP produces captions describing the highlighted object; the same interface supports region-focused VQA.

    For quantitative region-level captioning, the authors fine-tune LLaVA with the Alpha-CLIP image encoder frozen and evaluate on RefCOCOg and Visual Genome. The reported METEOR and CIDEr scores are:

    Could not parse LaTeX table

    Alpha-CLIP+LLaVA achieves the highest reported score in every metric shown for these two datasets.

  9. Knowl 9 — Region-controlled 2D image variation

    experimental setup

    The authors replace the ViT-L/14 CLIP image encoder in BLIP-Diffusion with Alpha-CLIP while leaving the remaining diffusion system unchanged. They use an empty text prompt so that the experiment tests the effect of the visual region prompt rather than text semantics. An alpha map selects the subject or image area whose appearance should guide variation, enabling subject-driven generation from complex scenes rather than requiring one centered foreground object.

    The qualitative comparison on page 7 evaluates Alpha-CLIP against image cropping, pixel-level image masking, a red-circle prompt, and feature-level masking. Cropping loses context and cannot reliably handle occlusion; a red circle changes the image content; pixel- and feature-level masking omit background information. Alpha-CLIP produces cleaner region-focused variations while retaining the original background and contextual relationships.

  10. Knowl 10 — Alpha-CLIP-guided Point-E point-cloud generation

    empirical result

    For image-to-3D generation, the authors replace the ViT-L/14 image encoder in the Point-E base-40M model with Alpha-CLIP and retain the rest of Point-E. A user-specified alpha map provides additional control over which parts of the input image should influence the generated point cloud. In the qualitative results on page 8, highlighting a missing part of the source object helps Point-E recover that part, while highlighting an important region causes Point-E to allocate more of its fixed budget of 1,024 points to that region. Thus, Alpha-CLIP adds region-level control to generalized image-to-3D generation without modifying Point-E’s diffusion model.

  11. Knowl 11 — Alpha-CLIP-guided PureCLIPNeRF optimization

    empirical result

    For text-to-3D optimization, the authors replace CLIP with Alpha-CLIP in PureCLIPNeRF. Each rendered NeRF image is supplied with an alpha channel obtained from NeRF density integration, allowing optimization gradients to pass through the alpha channel and directly influence the density parameters associated with the foreground object. Compared with the original CLIP, the Alpha-CLIP version produces objects whose shapes and colors more closely follow the text prompt, with improved global coherence and aesthetic quality in the qualitative results on page 8.

    The authors also test removing PureCLIPNeRF’s background-augmentation procedure. In most examples, Alpha-CLIP without background augmentation produces clearer, better text-aligned objects than original CLIP with augmentation, while running at approximately twice the speed. This shows that region-aware supervision can partly replace the time-consuming need to process multiple background-augmented views.

Coverage note — No substantial contributed material was omitted; supplementary ablations and additional qualitative visualizations were not made separate knowls because they support the architecture, training, and downstream results already covered.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736, 2022.
  2. 2.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  3. 3.Jinze Bai, Rui Men, Hao Yang, Xuancheng Ren, Kai Dang, Yichang Zhang, Xiaohuan Zhou, Peng Wang, Sinan Tan, An Yang, et al. Ofasys: A multi-modal multi-task learning system for building generalist models. arXiv preprint arXiv:2212.04408, 2022.
  4. 4.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023.
  5. 5.Keyan Chen, Xiaolong Jiang, Yao Hu, Xu Tang, Yan Gao, Jianqi Chen, and Weidi Xie. Ovarnet: Towards open-vocabulary object attribute recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23518–23527, 2023.
  6. 6.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  7. 7.Zheng Ding, Jie Wang, and Zhuowen Tu. Open-vocabulary universal image segmentation with maskclip. In International Conference on Machine Learning, 2022.
  8. 8.Xiaoyi Dong, Jianmin Bao, Yinglin Zheng, Ting Zhang, Dongdong Chen, Hao Yang, Ming Zeng, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, and Nenghai Yu. Maskclip: Masked self-distillation advances contrastive language-image pretraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10995–11005, 2023.
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  10. 10.Shanghua Gao, Zhong-Yu Li, Ming-Hsuan Yang, Ming-Ming Cheng, Junwei Han, and Philip Torr. Large-scale unsupervised semantic segmentation. IEEE transactions on pattern analysis and machine intelligence, 2022.
  11. 11.Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open-vocabulary object detection via vision and language knowledge distillation. In International Conference on Learning Representations, 2021.
  12. 12.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5356–5364, 2019.
  13. 13.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017.
  14. 14.MD Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga. A comprehensive survey of deep learning for image captioning. ACM Computing Surveys (CsUR), 51(6):1–36, 2019.
  15. 15.Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, et al. Language is not all you need: Aligning perception with language models. arXiv preprint arXiv:2302.14045, 2023.
  16. 16.Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hanneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Openclip. https : / / github . com / mlfoundations / open_clip, 2021.
  17. 17.Ajay Jain, Ben Mildenhall, Jonathan T Barron, Pieter Abbeel, and Ben Poole. Zero-shot text-guided object generation with dream fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 867–876, 2022.
  18. 18.John R Kender, Parijat Dube, Zhengyang Han, and Bishwaranjan Bhattacharjee. G2l: A high-dimensional geometric approach for automatic generation of highly accurate pseudo-labels. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1093–1102, 2023.
  19. 19.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023.
  20. 20.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalandidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123:32–73, 2017.
  21. 21.S Chandeesh Kumar, M Hemalatha, S Badri Narayan, and P Nandhini. Region driven remote sensing image captioning. Procedia Computer Science, 165:32–40, 2019.
  22. 22.Han-Hung Lee and Angel X Chang. Understanding pure clip guidance for voxel grid nerf models. arXiv preprint arXiv:2209.15172, 2022.
  23. 23.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. Otter: A multi-modal model with in-context instruction tuning. arXiv preprint arXiv:2305.03726, 2023.
  24. 24.Dongxu Li, Junnan Li, and Steven CH Hoi. Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing. arXiv preprint arXiv:2305.14720, 2023.
  25. 25.Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, pages 19730–19742. PMLR, 2023.
  26. 26.Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. Grounded language-image pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10965–10975, 2022.
  27. 27.Yanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer, and Kaiming He. Scaling language-image pre-training via masking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 23390–23400, 2023.
  28. 28.Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu. Open-vocabulary semantic segmentation with mask-adapted clip. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7061–7070, 2023.
  29. 29.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744, 2023.
  30. 30.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023.
  31. 31.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 11–20, 2016.
  32. 32.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021.
  33. 33.Nasir Mohammad Khalid, Tianhao Xie, Eugene Belilovsky, and Tiberiu Popa. Clip-mesh: Generating textured meshes from text using pretrained image-text models. In SIGGRAPH Asia 2022 conference papers, pages 1–8, 2022.
  34. 34.Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-e: A system for generating 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751, 2022.
  35. 35.OpenAI. Gpt-4 technical report, 2023.
  36. 36.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023.
  37. 37.Yanyuan Qiao, Chaorui Deng, and Qi Wu. Referring expression comprehension: A survey of methods and datasets. IEEE Transactions on Multimedia, 23:4426–4440, 2020.
  38. 38.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  39. 39.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  40. 40.Hanoona Rasheed, Muhammad Maaz, Sahal Shaji, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M. Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S. Khan. Glamm: Pixel grounding large multimodal model, 2023.
  41. 41.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022.
  42. 42.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22500–22510, 2023.
  43. 43.Aditya Sanghi, Hang Chu, Joseph G Lambourne, Ye Wang, Chin-Yi Cheng, Marco Fumero, and Kamal Rahimi Malekshan. Clip-forge: Towards zero-shot text-to-shape generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18603–18613, 2022.
  44. 44.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021.
  45. 45.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35:25278–25294, 2022.
  46. 46.Jing Shi, Wei Xiong, Zhe Lin, and Hyun Joon Jung. Instant-booth: Personalized text-to-image generation without test-time finetuning. arXiv preprint arXiv:2304.03411, 2023.
  47. 47.Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi. What does clip know about a red circle? visual prompt engineering for vlms. arXiv preprint arXiv:2304.06712, 2023.
  48. 48.Wei Su, Peihan Miao, Huanzhang Dou, Yongjian Fu, and Xi Li. Referring expression comprehension using language adaptive inference. arXiv preprint arXiv:2306.04451, 2023.
  49. 49.Sanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner, Sameer Singh, and Anna Rohrbach. Reclip: A strong zero-shot baseline for referring expression comprehension. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5198–5215, 2022.
  50. 50.Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. EVA-CLIP: improved training techniques for CLIP at scale. CoRR, abs/2303.15389, 2023.
  51. 51.Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative pretraining in multi-modality. arXiv preprint arXiv:2307.05222, 2023.
  52. 52.Vicuna. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. https://vicuna.lmsys.org/, 2023.
  53. 53.Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, and Thomas Wolf. Diffusers: State-of-the-art diffusion models. https://github.com/huggingface/diffusers, 2022.
  54. 54.Weiyun Wang, Min Shi, Qingyun Li, Wenhai Wang, Zhenhang Huang, Linjie Xing, Zhe Chen, Hao Li, Xizhou Zhu, Zhiguo Cao, Yushi Chen, Tong Lu, Jifeng Dai, and Yu Qiao. The all-seeing project: Towards panoptic visual recognition and understanding of the open world, 2023.
  55. 55.Yuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang, and Wangmeng Zuo. ELITE: Encoding visual concepts into textual embeddings for customized text-to-image generation. arXiv preprint arXiv:2302.13848, 2023.
  56. 56.Jialian Wu, Jianfeng Wang, Zhengyuan Yang, Zhe Gan, Zicheng Liu, Junsong Yuan, and Lijuan Wang. Grit: A generative region-to-text transformer for object understanding. arXiv preprint arXiv:2212.00280, 2022.
  57. 57.Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xiaolong Wang, and Shalini De Mello. Open-vocabulary panoptic segmentation with text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2955–2966, 2023.
  58. 58.Jiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao, Ying Shan, Xiaohu Qie, and Shenghua Gao. Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20908–20918, 2023.
  59. 59.Mengde Xu, Zheng Zhang, Fangyun Wei, Han Hu, and Xiang Bai. Side adapter network for open-vocabulary semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2945–2954, 2023.
  60. 60.Xin Xu, Tianyi Xiong, Zheng Ding, and Zhuowen Tu. Masq-clip for open-vocabulary universal image segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 887–898, 2023.
  61. 61.Lingfeng Yang, Yueze Wang, Xiang Li, Xinlong Wang, and Jian Yang. Fine-grained visual prompting. arXiv preprint arXiv:2306.04356, 2023.
  62. 62.Yuan Yao, Ao Zhang, Zhengyan Zhang, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun. Cpt: Colorful prompt tuning for pre-trained vision-language models. arXiv preprint arXiv:2109.11797, 2021.
  63. 63.Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv preprint arXiv:2308.06721, 2023.
  64. 64.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling context in referring expressions. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14, pages 69–85. Springer, 2016.
  65. 65.Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. Mattnet: Modular attention network for referring expression comprehension. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1307–1315, 2018.
  66. 66.Pan Zhang, Xiaoyi Dong Bin Wang, Yuhang Cao, Chao Xu, Linke Ouyang, Zhiyuan Zhao, Shuangrui Ding, Songyang Zhang, Haodong Duan, Hang Yan, et al. Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv preprint arXiv:2309.15112, 2023.
  67. 67.Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Kai Chen, and Ping Luo. Gpt4roi: Instruction tuning large language model on region-of-interest. arXiv preprint arXiv:2307.03601, 2023.
  68. 68.Shiyu Zhao, Zhixing Zhang, Samuel Schulter, Long Zhao, BG Vijay Kumar, Anastasis Stathopoulos, Manmohan Chandraker, and Dimitris N Metaxas. Exploiting unlabeled data with vision and language models for object detection. In European Conference on Computer Vision, pages 159–175. Springer, 2022.
  69. 69.Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, et al. Regionclip: Region-based language-image pretraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16793–16803, 2022.
  70. 70.Chong Zhou, Chen Change Loy, and Bo Dai. Extract free dense labels from clip. In European Conference on Computer Vision, pages 696–712. Springer, 2022.
  71. 71.Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krahenbühl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. In European Conference on Computer Vision, pages 350–368. Springer, 2022.
  72. 72.Yijie Zhou, Likun Cai, Xianhui Cheng, Zhongxue Gan, Xiangyang Xue, and Wenchao Ding. Openannotate3d: Open-vocabulary auto-labeling system for multi-modal 3d data. arXiv preprint arXiv:2310.13398, 2023.

Citation

MLA
Sun, Z., et al. “Alpha-CLIP: A CLIP Model Focusing on Wherever You Want”. arXiv, 2023, http://arxiv.org/abs/2312.03818v2.
APA
Sun, Z., Fang, Y., Wu, T., Zhang, P., Zang, Y., Kong, S., Xiong, Y., Lin, D., & Wang, J. (2023). Alpha-CLIP: A CLIP Model Focusing on Wherever You Want. arXiv. http://arxiv.org/abs/2312.03818v2
Chicago
Sun, Z., Y. Fang, T. Wu, et al. 2023. “Alpha-CLIP: A CLIP Model Focusing on Wherever You Want”. arXiv. http://arxiv.org/abs/2312.03818v2.
Harvard
Sun, Z. et al. (2023) “Alpha-CLIP: A CLIP Model Focusing on Wherever You Want”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.03818v2.
Vancouver
1. Sun Z, Fang Y, Wu T, Zhang P, Zang Y, Kong S, Xiong Y, Lin D, Wang J (2023) Alpha-CLIP: A CLIP Model Focusing on Wherever You Want. arXiv

BibTeX

@article{sun2023alpha,
  title = {Alpha-CLIP: A CLIP Model Focusing on Wherever You Want},
  author = {Sun, Zeyi and Fang, Ye and Wu, Tong and Zhang, Pan and Zang, Yuhang and Kong, Shu and Xiong, Yuanjun and Lin, Dahua and Wang, Jiaqi},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.03818v2},
  eprint = {2312.03818}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE