FG-CLIP: Fine-Grained Visual and Textual Alignment
Chunyu XieBin WangFanjing KongJincheng LiDawei LiangGengshen ZhangDawei LengYuhui Yin
Proposes FG-CLIP, a visual-language model that overcomes CLIP's coarse-grained limitations by training on 1.6 billion long captions, 40 million region-specific descriptions, and 10 million hard negative samples to boost performance in open-vocabulary object detection and image-text retrieval.
Modern vision-language artificial intelligence models, such as standard Contrastive Language-Image Pre-training (CLIP), have demonstrated strong capabilities in linking full images with short text descriptions. However, existing systems frequently fail when tasked with fine-grained visual understanding, such as distinguishing specific object attributes or identifying local regions within a complex scene. This limitation stems primarily from models being trained on brief text captions (often restricted to 77 tokens), relying strictly on whole-image matching without region-level context, and lacking challenging negative examples during training to help differentiate subtle visual differences.
The main objective of the article is to develop and evaluate Fine-Grained CLIP (FG-CLIP), a model architecture and two-stage training methodology designed to enhance fine-grained visual and textual alignment across global and local image features.
To achieve this, the authors constructed large-scale, high-quality multimodal datasets and implemented a dual-stage training framework. In the first stage, the model performs global contrastive learning using 1.6 billion image pairs paired with both short and detailed long captions (extended to support up to 248 tokens) generated by multimodal models. In the second stage, the model trains on a newly created dataset named FineHARD, consisting of 12 million images, 40 million region-specific bounding boxes with detailed descriptions, and 10 million automated "hard" negative samples where specific descriptive attributes were modified. The authors evaluated the system across standardized benchmarks for fine-grained understanding, bounding box classification, open-vocabulary object detection, image-text retrieval, and multimodal reasoning.
The evaluation yielded several key findings. First, FG-CLIP achieved dramatic gains in fine-grained understanding on the benchmark tests, scoring 48.4% accuracy on the hardest subset compared to only 15.4% for baseline CLIP and 22.8% for prior leading specialized models. Second, the model substantially improved localized region classification; on the LVIS dataset, bounding box classification accuracy rose to 38.3% compared to baseline CLIP's 9.3%. Third, FG-CLIP demonstrated superior cross-modal retrieval across both short and long text benchmarks, achieving 97.4% image-to-text retrieval on detailed captions where baseline CLIP achieved 86.5%. Finally, when integrated as the visual backbone for advanced vision-language systems like LLaVA, FG-CLIP improved object localization and attribute-based question answering while reducing output hallucinations.
These findings indicate that pairing large-scale detailed recaptioning with region-level alignment and hard negative samples resolves key structural deficiencies in multimodal AI. For operational applications, this translates directly to higher accuracy and reduced visual errors in automated surveillance, visual inspection, retrieval engines, and complex robotic interaction without requiring manual annotation at scale.
For organizations developing or deploying visual AI systems, the authors recommend adopting extended token lengths, incorporating region-level grounding into training pipelines, and leveraging automated hard-negative generation. Practitioners should also consider using FG-CLIP or its public dataset (FineHARD) as a drop-in replacement for standard visual backbones to improve downstream localization and reduce hallucinations.
Regarding limitations, curating captions and region-specific boxes relies partly on automated machine generation, which carries a minor noise rate (measured at approximately 1.1% in negative sample validation). In addition, generating and training across billions of multimodal pairs requires considerable computing infrastructure. Nevertheless, because the empirical improvements remain consistent across diverse public benchmarks and ablation tests, confidence in the core performance gains remains high.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. Read the original CLIP work first to understand the global image–text contrastive objective that FG-CLIP extends with longer captions and local alignment.
- Paper: RegionCLIP: Region-based Language-Image Pretraining, Yiwu Zhong et al. (2022). RegionCLIP establishes the region-level language alignment and automated region supervision that provide direct context for FG-CLIP’s localized training stage.
- Paper: Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts, Yan Zeng et al. (2022). X-VLM introduces multi-grained alignment across objects, regions, and whole images, clarifying the granularity problem FG-CLIP addresses.
- Paper: Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning, Jishnu Mukhoti et al. (2023). PACL shows how patch-level visual features can be aligned with text, preparing readers for FG-CLIP’s emphasis on local visual–textual correspondence.
- Paper: Alpha-CLIP: A CLIP Model Focusing on Wherever you Want, Zeyi Sun et al. (2024). Alpha-CLIP demonstrates adapting CLIP to attend to specified image regions, offering useful context for FG-CLIP’s region-grounded visual representations.
No sufficiently relevant recommendations were found.
