Image Difference Captioning with Pre-training and Contrastive Learning
Linli YaoWeiying WangQin Jin
Proposes a self-supervised pre-training and contrastive learning framework paired with a cross-task data expansion strategy to achieve state-of-the-art fine-grained image difference captioning despite limited annotated pairs.
Automated image difference captioning enables computer systems to compare two similar images and describe their visual distinctions in natural language. This capability is critical for practical applications such as medical lesion detection, surveillance monitoring, and fine-grained species identification. However, the task faces two major hurdles: models struggle to associate subtle visual differences with precise language, and creating human-annotated image pairs paired with descriptive text is expensive, leaving available training datasets relatively small.
The article demonstrates a new machine learning framework designed to improve image difference captioning by aligning visual and textual details through pre-training and contrastive learning. It also evaluates a data expansion strategy that incorporates broader image datasets to overcome the shortage of specialized difference-captioning data.
The approach uses a two-stage training strategy comprising self-supervised pre-training followed by task-specific fine-tuning. The architecture combines an image difference encoder, which locates distinctions across image pairs, with a cross-modal transformer that connects visual features to language. During pre-training, the model learns through three self-supervised objectives: masked language modeling, masked visual contrastive learning, and fine-grained difference alignment using strategically altered text samples. To supplement training, the framework integrates external datasets from general image captioning and fine-grained visual classification. The model was evaluated on two distinct benchmarks: CLEVR-Change, a synthetic dataset of geometric scene changes, and Birds-to-Words, a real-world dataset describing subtle differences between bird species.
The analysis produced several key findings. First, the proposed framework set new performance benchmarks on both datasets, increasing the primary quality score on CLEVR-Change from 118.7 to 128.9 and on Birds-to-Words from 45.6 to 48.4 without external data. Second, adding external visual and caption datasets further raised performance on Birds-to-Words, achieving a primary score of 49.1. Third, ablation testing showed that visual contrastive learning and fine-grained alignment were the largest contributors to model accuracy, while omitting the image difference encoder severely reduced performance. Finally, adding a specific classification task to filter out non-semantic distractions, such as shifts in camera angle or lighting, markedly improved performance on the synthetic benchmark.
These findings indicate that self-supervised pre-training and contrastive alignment allow AI systems to reliably identify and describe nuanced visual changes without requiring massive, costly human-annotated difference datasets. By leveraging existing single-image and classification data, organizations can lower data labeling costs and deployment risks for visual monitoring applications, improving reliability across synthetic and natural environments.
Organizations developing visual inspection or difference-reporting systems should adopt pre-training and contrastive learning strategies rather than relying solely on direct supervised training. When domain-specific difference data is scarce, teams should supplement training pipelines with relevant single-image classification and captioning datasets to build background knowledge.
The primary limitations noted in the article include a tendency of the model to occasionally repeat described differences or overlook semantic contradictions within generated sentences. While confidence in the benchmark results is high across the tested domains, practitioners should conduct pilot testing before deploying the framework in safety-critical operational workflows.
- Paper: Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, Junnan Li et al. (2021). Introduces the align-before-fuse contrastive pre-training paradigm that underpins modern fine-grained cross-modal alignment methods.
- Paper: Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks, Xiujun Li et al. (2020). Establishes object-tag semantic anchoring to ground ambiguous visual regions in language during multimodal pre-training.
- Paper: UNITER: UNiversal Image-TExt Representation Learning, Yen-Chun Chen et al. (2020). Provides the foundational multi-task pre-training objectives for aligning fine-grained visual regions with textual tokens.
- Paper: ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks, Jiasen Lu et al. (2019). Pioneers two-stream co-attentional Transformer pre-training for aligning visual regions with language representations.
- Paper: Deep visual-semantic alignments for generating image descriptions, Andrej Karpathy et al. (2015). Formulates the foundational visual-semantic alignment and caption generation framework for connecting image regions to descriptions.
- Paper: Momentum Contrast for Unsupervised Visual Representation Learning, Kaiming He et al. (2020). Introduces momentum contrastive learning, establishing core mechanisms for self-supervised representation learning.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). Surveys the broader ecosystem of vision-language architectures, pre-training objectives, and downstream adaptation strategies.
- Paper: CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection, Chuofan Ma et al. (2023). Extends fine-grained region-word visual-language alignment to open-vocabulary object discovery and localization.
- Paper: Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning, Jishnu Mukhoti et al. (2023). Applies localized contrastive alignment principles to achieve patch-level open-vocabulary semantic segmentation.
- Paper: Improving CLIP Training with Language Rewrites, Lijie Fan et al. (2023). Explores language rewriting strategies to improve text diversity and mitigate overfitting in contrastive vision-language pre-training.
- Paper: Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID, Wentao Tan et al. (2024). Applies synthetic caption generation and cross-modal alignment techniques to overcome data scarcity in fine-grained retrieval tasks.
