FG-CLIP: Fine-Grained Visual and Textual Alignment

Chunyu XieBin WangFanjing KongJincheng LiDawei LiangGengshen ZhangDawei LengYuhui Yin

article2025ICML114 citations

Proposes FG-CLIP, a visual-language model that overcomes CLIP's coarse-grained limitations by training on 1.6 billion long captions, 40 million region-specific descriptions, and 10 million hard negative samples to boost performance in open-vocabulary object detection and image-text retrieval.

Listen

Modern vision-language artificial intelligence models, such as standard Contrastive Language-Image Pre-training (CLIP), have demonstrated strong capabilities in linking full images with short text descriptions. However, existing systems frequently fail when tasked with fine-grained visual understanding, such as distinguishing specific object attributes or identifying local regions within a complex scene. This limitation stems primarily from models being trained on brief text captions (often restricted to 77 tokens), relying strictly on whole-image matching without region-level context, and lacking challenging negative examples during training to help differentiate subtle visual differences.

The main objective of the article is to develop and evaluate Fine-Grained CLIP (FG-CLIP), a model architecture and two-stage training methodology designed to enhance fine-grained visual and textual alignment across global and local image features.

To achieve this, the authors constructed large-scale, high-quality multimodal datasets and implemented a dual-stage training framework. In the first stage, the model performs global contrastive learning using 1.6 billion image pairs paired with both short and detailed long captions (extended to support up to 248 tokens) generated by multimodal models. In the second stage, the model trains on a newly created dataset named FineHARD, consisting of 12 million images, 40 million region-specific bounding boxes with detailed descriptions, and 10 million automated "hard" negative samples where specific descriptive attributes were modified. The authors evaluated the system across standardized benchmarks for fine-grained understanding, bounding box classification, open-vocabulary object detection, image-text retrieval, and multimodal reasoning.

The evaluation yielded several key findings. First, FG-CLIP achieved dramatic gains in fine-grained understanding on the benchmark tests, scoring 48.4% accuracy on the hardest subset compared to only 15.4% for baseline CLIP and 22.8% for prior leading specialized models. Second, the model substantially improved localized region classification; on the LVIS dataset, bounding box classification accuracy rose to 38.3% compared to baseline CLIP's 9.3%. Third, FG-CLIP demonstrated superior cross-modal retrieval across both short and long text benchmarks, achieving 97.4% image-to-text retrieval on detailed captions where baseline CLIP achieved 86.5%. Finally, when integrated as the visual backbone for advanced vision-language systems like LLaVA, FG-CLIP improved object localization and attribute-based question answering while reducing output hallucinations.

These findings indicate that pairing large-scale detailed recaptioning with region-level alignment and hard negative samples resolves key structural deficiencies in multimodal AI. For operational applications, this translates directly to higher accuracy and reduced visual errors in automated surveillance, visual inspection, retrieval engines, and complex robotic interaction without requiring manual annotation at scale.

For organizations developing or deploying visual AI systems, the authors recommend adopting extended token lengths, incorporating region-level grounding into training pipelines, and leveraging automated hard-negative generation. Practitioners should also consider using FG-CLIP or its public dataset (FineHARD) as a drop-in replacement for standard visual backbones to improve downstream localization and reduce hallucinations.

Regarding limitations, curating captions and region-specific boxes relies partly on automated machine generation, which carries a minor noise rate (measured at approximately 1.1% in negative sample validation). In addition, generating and training across billions of multimodal pairs requires considerable computing infrastructure. Nevertheless, because the empirical improvements remain consistent across diverse public benchmarks and ablation tests, confidence in the core performance gains remains high.

Xie et al (2025).pdf

No sufficiently relevant recommendations were found.

Cover for FG-CLIP: Fine-Grained Visual and Textual Alignment

Abstract

Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short captions. To address this, we propose Fine-Grained CLIP (FG-CLIP), which enhances fine-grained understanding through three key innovations. First, we leverage large multimodal models to generate 1.6 billion long caption-image pairs for capturing global-level semantic details. Second, a high-quality dataset is constructed with 12 million images and 40 million region-specific bounding boxes aligned with detailed captions to ensure precise, context-rich representations. Third, 10 million hard fine-grained negative samples are incorporated to improve the model’s ability to distinguish subtle semantic differences. We construct a comprehensive dataset, termed FineHARD, by integrating high-quality region-specific annotations with hard fine-grained negative samples. Corresponding training methods are meticulously designed for these data. Extensive experiments demonstrate that FG-CLIP outperforms the original CLIP and other state-of-the-art methods across various downstream tasks, including fine-grained understanding, open-vocabulary object detection, image-text retrieval, and general multimodal benchmarks. These results highlight FG-CLIP’s effectiveness in capturing fine-grained image details and improving overall model performance. The data, code, and models are available at https://github.com/360CVGroup/FG-CLIP.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Contrastive Language-Image Pre-training
  • 2.2. Fine-Grained Understanding
  • 2.3. Image-Text Datasets
  • 3. Approach
  • 3.1. Fine-Grained CLIP
  • 3.2. Curated Dataset
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Comparisons on Fine-grained Region-level Task
  • 4.3. Comparisons on Image-level Task
  • 4.4. Comparisons on General Multimodal Benchmarks
  • 4.5. Ablation Study
  • 5. Conclusion
  • Impact Statement
  • References
  • A. Examples of Curated Visual Grounding Data
  • B. Positive and Negative Descriptions Related to Image Regions
  • C. Visualization Comparison
  • D. Further Experiments
  • D.1. Comparison of Different Methods on Fine-Grained Benchmark
  • D.2. Performance Comparison on Identical Datasets
  • D.3. Performance on General Multimodal Benchmarks

Knowls

  1. Knowl 1 — Two-stage training separates global and region-level alignment

    model/method

    FG-CLIP retains CLIP’s dual-encoder structure and trains it in two stages. Stage 1 uses global image–text contrastive learning with short and detailed captions. Stage 2 starts from the Stage 1 model and combines global contrastive learning with region–text alignment and hard-negative learning. Its objective is

    L=Lglobal+αLregional+βLhard,L = L_{\mathrm{global}} + \alpha L_{\mathrm{regional}} + \beta L_{\mathrm{hard}},

    where the regional and hard-negative weights are α=0.1\alpha=0.1 and β=0.5\beta=0.5. Stage 1 is initialized from original CLIP weights; each stage is trained for one epoch.

  2. Knowl 2 — Detailed captions extend global image–text alignment

    model/method

    For global alignment, FG-CLIP retains short captions and adds detailed long captions generated for 1.6 billion image examples using CogVLM2-19B. Both caption types are aligned to the image using the class-token embeddings from the text and image encoders. The text encoder’s positional embeddings are adapted for longer inputs: sequences of at most 20 tokens use the original embeddings, while positions beyond 20 use linear interpolation with a factor of 4. This extends the maximum supported sequence length from 77 to 248 tokens.

    For a batch of NN matched image–caption pairs, let vi∈Rdv_i\in\mathbb{R}^d and ti∈Rdt_i\in\mathbb{R}^d be the image and text embeddings, respectively, where dd is the embedding dimension. The cosine similarity is s(v,t)=v⊤t/(∥v∥∥t∥)s(v,t)=v^\top t/(\|v\|\|t\|). The symmetric global contrastive loss is

    Lglobal=−12N∑i=1N[log⁡exp⁡(s(vi,ti)/τ)∑j=1Nexp⁡(s(vi,tj)/τ)+log⁡exp⁡(s(ti,vi)/τ)∑j=1Nexp⁡(s(ti,vj)/τ)],L_{\mathrm{global}}=-\frac{1}{2N}\sum_{i=1}^{N}\left[\log\frac{\exp(s(v_i,t_i)/\tau)}{\sum_{j=1}^{N}\exp(s(v_i,t_j)/\tau)}+\log\frac{\exp(s(t_i,v_i)/\tau)}{\sum_{j=1}^{N}\exp(s(t_i,v_j)/\tau)}\right],

    where τ\tau is a learnable temperature. The objective is applied to global image–text matching; the training examples include both short and long captions.

  3. Knowl 3 — FineHARD supplies region annotations and hard negatives at scale

    data/table

    FineHARD is the second-stage dataset, curated from images in GRIT. CogVLM2-19B generates detailed image captions; spaCy extracts referring expressions, and YOLO-World predicts their bounding boxes. Non-maximum suppression removes overlapping boxes, and only detections with confidence greater than 0.4 are retained. The resulting corpus contains 12 million images and 40 million bounding boxes with region-specific descriptions.

    To make hard negatives, Llama-3.1-70B rewrites bounding-box descriptions by changing attributes while keeping object names unchanged. The process generates 10 negative descriptions per positive description and yields 10 million hard-negative samples overall. In a quality check of 3,000 negative samples, 98.9% were judged qualified and 1.1% were considered noise. The data-processing pipeline used a cluster of 160×910B NPUs and took seven days.

  4. Knowl 4 — Regional contrastive learning aligns image regions with text segments

    equation

    FG-CLIP extracts region features with RoIAlign and average-pools the visual tokens inside each bounding box. It matches each resulting region embedding to a corresponding phrase or sentence from the image caption. For a batch containing KK valid region–text matches, let rir_i and lil_i be the visual-region and text embeddings for match ii, and let s(a,b)=a⊤b/(∥a∥∥b∥)s(a,b)=a^\top b/(\|a\|\|b\|) be cosine similarity. The symmetric regional contrastive loss is

    Lregional=−12K∑i=1K[log⁡exp⁡(s(ri,li)/τ)∑j=1Kexp⁡(s(ri,lj)/τ)+log⁡exp⁡(s(li,ri)/τ)∑j=1Kexp⁡(s(li,rj)/τ)],L_{\mathrm{regional}}=-\frac{1}{2K}\sum_{i=1}^{K}\left[\log\frac{\exp(s(r_i,l_i)/\tau)}{\sum_{j=1}^{K}\exp(s(r_i,l_j)/\tau)}+\log\frac{\exp(s(l_i,r_i)/\tau)}{\sum_{j=1}^{K}\exp(s(l_i,r_j)/\tau)}\right],

    where τ\tau is the learnable contrastive temperature. The denominators use the other region or text embeddings in the batch as mismatched candidates.

  5. Knowl 5 — Hard-negative learning distinguishes subtle region attributes

    equation

    FG-CLIP uses hard negatives that describe semantically similar but non-matching regions. For each region, the negative descriptions are produced by changing attributes in its positive description while leaving the object name unchanged. Let rir_i be the visual embedding of region ii, and let li,1l_{i,1} be its positive text embedding while li,jl_{i,j} for j>1j>1 are negative text embeddings. With KK regions, MM candidate descriptions per region, cosine similarity s(a,b)=a⊤b/(∥a∥∥b∥)s(a,b)=a^\top b/(\|a\|\|b\|), and learnable temperature τ\tau, the hard-negative loss is

    Lhard=−1K∑i=1Klog⁡exp⁡(s(ri,li,1)/τ)∑j=1Mexp⁡(s(ri,li,j)/τ).L_{\mathrm{hard}}=-\frac{1}{K}\sum_{i=1}^{K}\log\frac{\exp(s(r_i,l_{i,1})/\tau)}{\sum_{j=1}^{M}\exp(s(r_i,l_{i,j})/\tau)}.

    This term raises the match score for the positive description relative to attribute-altered descriptions of the same region.

  6. Knowl 6 — Ablations isolate the effects of global, regional, and hard-negative learning

    empirical result

    An ablation evaluated long-caption retrieval on DCI (image-to-text/text-to-image), short-caption retrieval on MSCOCO (image-to-text/text-to-image), bounding-box classification on COCO (Top-1/Top-5), and FG-OVD fine-grained accuracy (hard/medium/easy). Values are reported in that order; retrieval and classification metrics are the benchmark scores reported by the paper, and FG-OVD values are percentages.

    Original CLIP scored 45.5/43.0, 51.8/32.7, 44.2/72.3, and 12.0/23.1/22.2. FG-CLIP Stage 1, trained with detailed captions, scored 58.3/57.5, 64.6/44.9, 47.2/74.2, and 21.8/41.6/36.2. Adding Stage 2 global contrastive learning gave 62.7/61.2, 64.4/46.4, 46.8/73.6, and 25.4/46.8/42.9. Adding regional learning gave 62.4/61.1, 64.7/45.7, 53.7/81.2, and 24.5/47.1/49.5. Adding hard-negative learning produced the full model’s scores: 61.8/60.6, 64.1/45.4, 52.3/79.7, and 46.1/66.6/68.7.

    The ablation shows that regional learning raises COCO box-classification Top-1 from 46.8 to 53.7, while hard-negative learning raises FG-OVD accuracy from 24.5 to 46.1 on hard examples, from 47.1 to 66.6 on medium examples, and from 49.5 to 68.7 on easy examples. Hard-negative learning therefore produces its largest measured gains on the fine-grained matching benchmark, with smaller changes in retrieval and box classification.

  7. Knowl 7 — FG-CLIP improves fine-grained region matching across difficulty levels

    empirical result

    On FG-OVD, models are evaluated by Top-1 accuracy (%) when matching a localized image region against one positive and ten negative descriptions. The benchmark’s hard, medium, and easy subsets replace one, two, and three attribute words, respectively; the trivial subset uses unrelated descriptions.

    For ViT-B/16, FG-CLIP achieved 46.1/66.6/68.7/83.4 on hard/medium/easy/trivial examples, compared with FineCLIP’s 26.8/49.8/50.4/71.9 and CLIP’s 12.0/23.1/22.2/58.5. For ViT-L/14, FG-CLIP achieved 48.4/69.5/71.2/89.7, compared with FineCLIP’s 22.8/46.0/46.0/73.6 and CLIP’s 15.4/25.3/25.7/38.8. FG-CLIP scored higher than these comparators in all four difficulty subsets for both backbones, with especially large gains on hard and medium cases.

  8. Knowl 8 — Localized zero-shot box classification improves on three datasets

    empirical result

    Zero-shot bounding-box classification was evaluated by Top-1 accuracy on COCO, LVIS, and Open Images, using provided boxes and text descriptions of the dataset categories. FG-CLIP ViT-B/16 scored 52.3, 28.6, and 20.6, respectively; the corresponding CLIP scores were 44.2, 20.9, and 15.3, and FineCLIP scores were 48.4, 23.3, and 18.1. FG-CLIP ViT-L/14 scored 63.2, 38.3, and 23.8, compared with CLIP ViT-L/14 at 33.8, 9.3, and 8.3, and FineCLIP ViT-L/14 at 54.5, 22.5, and 19.1. FG-CLIP attained the highest reported score among these methods on all three datasets at both backbone sizes.

  9. Knowl 9 — FG-CLIP improves novel-category detection as an F-ViT backbone

    empirical result

    The paper evaluated open-vocabulary detection on OV-COCO by using FG-CLIP in the F-ViT two-stage detector with a frozen visual encoder. The reported metric is box AP at IoU 0.5 for novel, base, and all categories, in that order. With ViT-B/16, F-ViT plus FG-CLIP scored 35.1/51.7/47.4, compared with 33.6/54.2/48.8 using CLIPSelf and 29.8/45.9/41.7 using FineCLIP. With ViT-L/14, F-ViT plus FG-CLIP scored 41.2/58.0/53.6, compared with 38.4/60.6/54.8 using CLIPSelf and 40.0/57.2/52.7 using FineCLIP. FG-CLIP gave the highest novel-category AP among these F-ViT variants at both scales; the highest base and all-category values in these comparisons were instead obtained with CLIPSelf.

  10. Knowl 10 — Image–text retrieval gains span long and short captions

    empirical result

    The paper evaluated long-caption retrieval on the 1K ShareGPT4V subset and the 7,805-pair DCI dataset, short-caption retrieval on MSCOCO 5K and Flickr 1K, and zero-shot classification on ImageNet-1K and ImageNet-v2. For each retrieval dataset, image-to-text and text-to-image scores are given in that order.

    FG-CLIP ViT-B/16 scored 96.7/94.9 on ShareGPT4V, 61.8/60.6 on DCI, 64.1/45.4 on MSCOCO, and 90.7/76.4 on Flickr; its ImageNet-1K and ImageNet-v2 Top-1 accuracies were 69.0 and 61.8. CLIP ViT-B/16 scored 78.2/79.6, 45.5/43.0, 51.8/32.7, and 82.2/62.1 on the four retrieval sets, and 68.4/61.9 on ImageNet-1K/ImageNet-v2. FG-CLIP ViT-L/14 scored 97.4/96.8, 66.7/66.1, 68.9/50.9, and 93.7/81.5 on retrieval, with classification accuracies of 76.1/69.0; CLIP ViT-L/14 scored 86.5/83.6, 37.2/36.4, 58.0/37.1, and 87.4/67.3, with classification accuracies of 76.6/70.9. FG-CLIP improves over CLIP on all reported retrieval measures, while its ImageNet classification scores are slightly below EVA-CLIP’s reported 74.7/67.0 for ViT-B/16 and 80.4/73.8 for ViT-L/14.

  11. Knowl 11 — Replacing CLIP with FG-CLIP improves LLaVA-v1.5 benchmark scores

    empirical result

    FG-CLIP was evaluated as the visual feature extractor in LLaVA-v1.5-7B, with the LLaVA parameter configuration and training data held consistent with the CLIP-based baseline. Scores for CLIP-based LLaVA versus FG-CLIP-based LLaVA were: GQA 61.9 versus 62.9; POPE 85.9 versus 86.8; RefCOCO validation 76.2 versus 81.4, testA 83.4 versus 86.5, and testB 67.9 versus 74.9; MMBench-EN development 65.1 versus 66.6 and test 66.5 versus 66.7; and MMBench-CN development 58.2 versus 58.8 and test 58.4 versus 59.3. The largest gains were on RefCOCO, which tests referring-expression localization and attribute understanding; gains on the other reported benchmarks were smaller.

Coverage note — The appendix’s qualitative similarity-map examples and detailed optimizer, batch-size, and hardware training settings are omitted: the visualizations corroborate the quantitative localization results, while the settings are implementation details rather than separate contributions.

References

  1. 1.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. In NeurIPS, volume 35, pp. 23716–23736, 2022.
  2. 2.Bianchi, L., Carrara, F., Messina, N., Gennaro, C., and Falchi, F. The devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understanding. In CVPR, pp. 22520–22529, 2024.
  3. 3.Changpinyo, S., Sharma, P., Ding, N., and Soricut, R. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR, pp. 3558–3568, 2021.
  4. 4.Chen, L., Li, J., Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., and Lin, D. Sharegpt4v: Improving large multi-modal models with better captions. In ECCV, pp. 370–387, 2024a.
  5. 5.Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In CVPR, pp. 24185–24198, 2024b.
  6. 6.Cheng, T., Song, L., Ge, Y., Liu, W., Wang, X., and Shan, Y. Yoloworld: Real-time open-vocabulary object detection. In CVPR, pp. 16901–16911, 2024.
  7. 7.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In CVPR, pp. 248–255, 2009.
  8. 8.Dosovitskiy, A. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  9. 9.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  10. 10.Fu, L., Datta, G., Huang, H., Panitch, W. C.-H., Drake, J., Ortiz, J., Mukadam, M., Lambeta, M., Calandra, R., and Goldberg, K. A touch, vision, and language dataset for multimodal alignment. In ICML, pp. 14080–14101, 2024.
  11. 11.Gabeff, V., Rußwurm, M., Tuia, D., and Mathis, A. Wildclip: Scene and animal attribute retrieval from camera trap data with domain-adapted vision-language models. IJCV, pp. 1–17, 2024.
  12. 12.Gu, J., Meng, X., Lu, G., Hou, L., Minzhe, N., Liang, X., Yao, L., Huang, R., Zhang, W., Jiang, X., et al. Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark. In NeurIPS, volume 35, pp. 26418–26431, 2022.
  13. 13.Gupta, A., Dollar, P., and Girshick, R. Lvis: A dataset for large vocabulary instance segmentation. In CVPR, pp. 5356–5364, 2019.
  14. 14.He, K., Gkioxari, G., Dollár, P., and Girshick, R. Mask r-cnn. In ICCV, pp. 2961–2969, 2017.
  15. 15.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum contrast for unsupervised visual representation learning. In CVPR, pp. 9729–9738, 2020.
  16. 16.Hong, W., Wang, W., Ding, M., Yu, W., Lv, Q., Wang, Y., Cheng, Y., Huang, S., Ji, J., Xue, Z., et al. Cogvlm2: Visual language models for image and video understanding. arXiv preprint arXiv:2408.16500, 2024.
  17. 17.Honnibal, M., Montani, I., Van Landeghem, S., Boyd, A., et al. spacy: Industrial-strength natural language processing in python. 2020.
  18. 18.Hou, X., Liu, M., Zhang, S., Wei, P., and Chen, B. Salience detr: Enhancing detection transformer with hierarchical salience filtering refinement. In CVPR, pp. 17574–17583, 2024.
  19. 19.Hudson, D. A. and Manning, C. D. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In CVPR, pp. 6700–6709, 2019.
  20. 20.Jing, D., He, X., Luo, Y., Fei, N., Yang, G., Wei, W., Zhao, H., and Lu, Z. Fineclip: Self-distilled region-based clip for better fine-grained understanding. In NeurIPS, 2024.
  21. 21.Kazemzadeh, S., Ordonez, V., Matten, M., and Berg, T. Referitgame: Referring to objects in photographs of natural scenes. In EMNLP, pp. 787–798, 2014.
  22. 22.Kim, D., Angelova, A., and Kuo, W. Contrastive feature masking open-vocabulary vision transformer. In ICCV, pp. 15602–15612, 2023a.
  23. 23.Kim, D., Angelova, A., and Kuo, W. Region-aware pretraining for open-vocabulary object detection with vision transformers. In CVPR, pp. 11144–11154, 2023b.
  24. 24.Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., et al. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. IJCV, 128(7):1956–1981, 2020.
  25. 25.Laurençon, H., Marafioti, A., Sanh, V., and Tronchon, L. Building and better understanding vision-language models: insights and future directions. arXiv preprint arXiv:2408.12637, 2024.
  26. 26.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, pp. 19730–19742, 2023a.
  27. 27.Li, J., Xie, C., Wu, X., Wang, B., and Leng, D. What makes good open-vocabulary detector: A disassembling perspective. arXiv preprint arXiv:2309.00227, 2023b.
  28. 28.Li, L. H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.-N., et al. Grounded language-image pre-training. In CVPR, pp. 10965–10975, 2022.
  29. 29.Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and Wen, J.-R. Evaluating object hallucination in large vision-language models. In EMNLP, 2023c. URL https://openreview.net/forum?id=xozJw0kZXF.
  30. 30.Li, Z., Yang, B., Liu, Q., Ma, Z., Zhang, S., Yang, J., Sun, Y., Liu, Y., and Bai, X. Monkey: Image resolution and text label are important things for large multi-modal models. In CVPR, 2024.
  31. 31.Lin, C., Sun, P., Jiang, Y., Luo, P., Qu, L., Haffari, G., Yuan, Z., and Cai, J. Learning object-language alignments for open-vocabulary object detection. arXiv preprint arXiv:2211.14843, 2022.
  32. 32.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In ECCV, pp. 740–755, 2014.
  33. 33.Lin, W., Zhao, Z., Zhang, X., Wu, C., Zhang, Y., Wang, Y., and Xie, W. Pmc-clip: Contrastive language-image pre-training using biomedical documents. In MICCAI, pp. 525–536, 2023.
  34. 34.Liu, C., Zhang, Y., Wang, H., Chen, W., Wang, F., Huang, Y., Shen, Y.-D., and Wang, L. Efficient token-guided image-text retrieval with consistent multimodal contrastive training. IEEE Transactions on Image Processing, 32:3622–3633, 2023a.
  35. 35.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. In NeurIPS, volume 36, pp. 34892–34916, 2023b.
  36. 36.Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., et al. Mmbench: Is your multi-modal model an all-around player? In ECCV, pp. 216–233, 2024.
  37. 37.Ma, C., Jiang, Y., Wu, J., Yuan, Z., and Qi, X. Groma: Localized visual tokenization for grounding multimodal large language models. In ECCV, pp. 417–435, 2024.
  38. 38.Minderer, M., Gritsenko, A., and Houlsby, N. Scaling open-vocabulary object detection. In NeurIPS, volume 36, 2024.
  39. 39.Mokady, R., Hertz, A., and Bermano, A. H. Clipcap: Clip prefix for image captioning. arXiv preprint arXiv:2111.09734, 2021.
  40. 40.Pan, J., Ma, Q., and Bai, C. A prior instruction representation framework for remote sensing image-text retrieval. In ACM MM, pp. 611–620, 2023.
  41. 41.Parelli, M., Delitzas, A., Hars, N., Vlassis, G., Anagnostidis, S., Bachmann, G., and Hofmann, T. Clip-guided vision-language pre-training for question answering in 3d scenes. In CVPR, pp. 5607–5612, 2023.
  42. 42.Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., Ye, Q., and Wei, F. Kosmos-2: Grounding multimodal large language models to the world. In ICLR, 2024.
  43. 43.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, A., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In ICML, pp. 8748–8763, 2021.
  44. 44.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1(2):3, 2022.
  45. 45.Recht, B., Roelofs, R., Schmidt, L., and Shankar, V. Do imagenet classifiers generalize to imagenet? In ICML, pp. 5389–5400, 2019.
  46. 46.Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021.
  47. 47.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. Laion-5b: An open large-scale dataset for training next generation image-text models. In NeurIPS, volume 35, pp. 25278–25294, 2022.
  48. 48.Sharma, P., Ding, N., Goodman, S., and Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL, pp. 2556–2565, 2018.
  49. 49.Sun, Q., Fang, Y., Wu, L., Wang, X., and Cao, Y. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023.
  50. 50.Sun, Z., Fang, Y., Wu, T., Zhang, P., Zang, Y., Kong, S., Xiong, Y., Lin, D., and Wang, J. Alpha-clip: A clip model focusing on wherever you want. In CVPR, pp. 13019–13029, 2024.
  51. 51.Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024.
  52. 52.Urbanek, J., Bordes, F., Astolfi, P., Williamson, M., Sharma, V., and Romero-Soriano, A. A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions. In CVPR, pp. 26700–26709, 2024.
  53. 53.Wang, B., Xie, C., Leng, D., and Yin, Y. Iaa: Inner-adaptor architecture empowers frozen large language model with multimodal capabilities. In AAAI, volume 39, pp. 21035–21043, 2025.
  54. 54.Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024.
  55. 55.Wu, S., Zhang, W., Xu, L., Jin, S., Li, X., Liu, W., and Loy, C. C. CLIPSelf: Vision transformer distills itself for open-vocabulary dense prediction. In ICLR, 2024a. URL https://openreview.net/forum?id=DjzvJCRsVf.
  56. 56.Wu, W., Zheng, K., Ma, S., Lu, F., Guo, Y., Zhang, Y., Chen, W., Guo, Q., Shen, Y., and Zha, Z.-J. Lotlip: Improving language-image pre-training for long text understanding. arXiv preprint arXiv:2410.05249, 2024b.
  57. 57.Wu, Z., Chen, X., Pan, Z., Liu, X., Liu, W., Dai, D., Gao, H., Ma, Y., Wu, C., Wang, B., et al. Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding. arXiv preprint arXiv:2412.10302, 2024c.
  58. 58.Xie, C., Li, C., Zhang, B., Han, J., Zhen, X., and Chen, J. Memory attention networks for skeleton-based action recognition. In IJCAI, pp. 1639–1645, 2018.
  59. 59.Xie, C., Cai, H., Li, J., Kong, F., Wu, X., Song, J., Morimitsu, H., Yao, L., Wang, D., Zhang, X., et al. Ccmb: A large-scale chinese cross-modal benchmark. In ACM MM, pp. 4219–4227, 2023.
  60. 60.Yao, Y., Yu, T., Zhang, A., Wang, C., Cui, J., Zhu, H., Cai, T., Li, H., Zhao, W., He, Z., et al. Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800, 2024.
  61. 61.Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78, 2014.
  62. 62.Zareian, A., Rosa, K. D., Hu, D. H., and Chang, S.-F. Open-vocabulary object detection using captions. In CVPR, pp. 14393–14402, 2021.
  63. 63.Zhang, B., Zhang, P., Dong, X., Zang, Y., and Wang, J. Long-clip: Unlocking the long-text capability of clip. In ECCV, pp. 310–325, 2024.
  64. 64.Zheng, K., Zhang, Y., Wu, W., Lu, F., Ma, S., Jin, X., Chen, W., and Shen, Y. Dreamlip: Language-image pre-training with long captions. In ECCV, pp. 73–90, 2024.
  65. 65.Zhong, Y., Yang, J., Zhang, P., Li, C., Codella, N., Li, L. H., Zhou, L., Dai, X., Yuan, L., Li, Y., et al. Regionclip: Region-based language-image pretraining. In CVPR, pp. 16793–16803, 2022.
  66. 66.Zhou, C., Loy, C. C., and Dai, B. Extract free dense labels from clip. In ECCV, pp. 696–712, 2022a.
  67. 67.Zhou, X., Girdhar, R., Joulin, A., Krähenbühl, P., and Misra, I. Detecting twenty-thousand classes using image-level supervision. In ECCV, pp. 350–368, 2022b.

Citation

MLA
Xie, C., et al. “FG-CLIP: Fine-Grained Visual and Textual Alignment”. arXiv, 2025, https://doi.org/10.48550/arxiv.2505.05071.
APA
Xie, C., Wang, B., Kong, F., Li, J., Liang, D., Zhang, G., Leng, D., & Yin, Y. (2025). FG-CLIP: Fine-Grained Visual and Textual Alignment. arXiv. https://doi.org/10.48550/arxiv.2505.05071
Chicago
Xie, C., B. Wang, F. Kong, et al. 2025. “FG-CLIP: Fine-Grained Visual and Textual Alignment”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2505.05071.
Harvard
Xie, C. et al. (2025) “FG-CLIP: Fine-Grained Visual and Textual Alignment”. arXiv. Available at: https://doi.org/10.48550/arxiv.2505.05071.
Vancouver
1. Xie C, Wang B, Kong F, Li J, Liang D, Zhang G, Leng D, Yin Y (2025) FG-CLIP: Fine-Grained Visual and Textual Alignment. https://doi.org/10.48550/arxiv.2505.05071

BibTeX

@misc{https://doi.org/10.48550/arxiv.2505.05071,
  doi = {10.48550/ARXIV.2505.05071},
  url = {https://arxiv.org/abs/2505.05071},
  author = {Xie, Chunyu and Wang, Bin and Kong, Fanjing and Li, Jincheng and Liang, Dawei and Zhang, Gengshen and Leng, Dawei and Yin, Yuhui},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {FG-CLIP: Fine-Grained Visual and Textual Alignment},
  publisher = {arXiv},
  year = {2025},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/