Image Difference Captioning with Pre-training and Contrastive Learning

Linli YaoWeiying WangQin Jin

article2022AAAI65 citations

Proposes a self-supervised pre-training and contrastive learning framework paired with a cross-task data expansion strategy to achieve state-of-the-art fine-grained image difference captioning despite limited annotated pairs.

Listen

Automated image difference captioning enables computer systems to compare two similar images and describe their visual distinctions in natural language. This capability is critical for practical applications such as medical lesion detection, surveillance monitoring, and fine-grained species identification. However, the task faces two major hurdles: models struggle to associate subtle visual differences with precise language, and creating human-annotated image pairs paired with descriptive text is expensive, leaving available training datasets relatively small.

The article demonstrates a new machine learning framework designed to improve image difference captioning by aligning visual and textual details through pre-training and contrastive learning. It also evaluates a data expansion strategy that incorporates broader image datasets to overcome the shortage of specialized difference-captioning data.

The approach uses a two-stage training strategy comprising self-supervised pre-training followed by task-specific fine-tuning. The architecture combines an image difference encoder, which locates distinctions across image pairs, with a cross-modal transformer that connects visual features to language. During pre-training, the model learns through three self-supervised objectives: masked language modeling, masked visual contrastive learning, and fine-grained difference alignment using strategically altered text samples. To supplement training, the framework integrates external datasets from general image captioning and fine-grained visual classification. The model was evaluated on two distinct benchmarks: CLEVR-Change, a synthetic dataset of geometric scene changes, and Birds-to-Words, a real-world dataset describing subtle differences between bird species.

The analysis produced several key findings. First, the proposed framework set new performance benchmarks on both datasets, increasing the primary quality score on CLEVR-Change from 118.7 to 128.9 and on Birds-to-Words from 45.6 to 48.4 without external data. Second, adding external visual and caption datasets further raised performance on Birds-to-Words, achieving a primary score of 49.1. Third, ablation testing showed that visual contrastive learning and fine-grained alignment were the largest contributors to model accuracy, while omitting the image difference encoder severely reduced performance. Finally, adding a specific classification task to filter out non-semantic distractions, such as shifts in camera angle or lighting, markedly improved performance on the synthetic benchmark.

These findings indicate that self-supervised pre-training and contrastive alignment allow AI systems to reliably identify and describe nuanced visual changes without requiring massive, costly human-annotated difference datasets. By leveraging existing single-image and classification data, organizations can lower data labeling costs and deployment risks for visual monitoring applications, improving reliability across synthetic and natural environments.

Organizations developing visual inspection or difference-reporting systems should adopt pre-training and contrastive learning strategies rather than relying solely on direct supervised training. When domain-specific difference data is scarce, teams should supplement training pipelines with relevant single-image classification and captioning datasets to build background knowledge.

The primary limitations noted in the article include a tendency of the model to occasionally repeat described differences or overlook semantic contradictions within generated sentences. While confidence in the benchmark results is high across the tested domains, practitioners should conduct pilot testing before deploying the framework in safety-critical operational workflows.

Cover for Image Difference Captioning with Pre-training and Contrastive Learning

Abstract

The Image Difference Captioning (IDC) task aims to describe the visual differences between two similar images with natural language. The major challenges of this task lie in two aspects: 1) fine-grained visual differences that require learning stronger vision and language association and 2) high-cost of manual annotations that leads to limited supervised data. To address these challenges, we propose a new modeling framework following the pre-training-finetuning paradigm. Specifically, we design three self-supervised tasks and contrastive learning strategies to align visual differences and text descriptions at a fine-grained level. Moreover, we propose a data expansion strategy to utilize extra cross-task supervision information, such as data for fine-grained image classification, to alleviate the limitation of available supervised IDC data. Extensive experiments on two IDC benchmark datasets, CLEVR-Change and Birds-to-Words, demonstrate the effectiveness of the proposed modeling framework. The codes and models will be released at https://github.com/yaolinli/IDC.

Table of Contents

  • Introduction
  • Method
  • Model Architecture
  • Pre-training Tasks
  • Finetuning and Inference
  • Expansion of Cross-task Data
  • Experiments
  • Experimental Settings
  • Comparison with the State-of-the-Arts
  • Results and Analysis
  • Related Works
  • Image Difference Captioning
  • Vision-Language Pre-training
  • Conclusions
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Architecture of the Image Difference Captioning Framework

    model/method

    The Image Difference Captioning (IDC) architecture models an input triplet comprising an image pair and a text description, denoted as {V(1),V(2),T}\{V^{(1)}, V^{(2)}, T\}, via a two-stage visual encoder followed by a multi-layer cross-modal Transformer.

    • Input Representation: The text description T={[CLS],[BOS],w0,…,wM,[EOS]}T = \{[\text{CLS}], [\text{BOS}], w_0, \dots, w_M, [\text{EOS}]\} is embedded into a 512-dimensional vector space learned from scratch, where [CLS][\text{CLS}] captures global sentence semantics. Each image is processed by a pre-trained ResNet-101 backbone to extract grid features of shape (49,2048)(49, 2048), projected via a linear layer to dimension 512, and prepended with a global image token: V(1)={[IMG1],v0(1),…,vN(1)}V^{(1)} = \{[\text{IMG1}], v_0^{(1)}, \dots, v_N^{(1)}\} and V(2)={[IMG2],v0(2),…,vN(2)}V^{(2)} = \{[\text{IMG2}], v_0^{(2)}, \dots, v_N^{(2)}\} (N=49N=49). Fixed sinusoidal positional embeddings and modality type embeddings are added to all visual and textual tokens.

    • Image Difference Encoder: Visual features pass through a shared single-image Transformer encoder FsingF_{\text{sing}} to capture intra-image regional semantics, followed by a pair-image Transformer encoder FpairF_{\text{pair}} that models cross-image interaction and locates differences:

    V~(1),V~(2)=Fpair(Fsing(V(1)),Fsing(V(2)))\tilde{V}^{(1)}, \tilde{V}^{(2)} = F_{\text{pair}}\left(F_{\text{sing}}(V^{(1)}), F_{\text{sing}}(V^{(2)})\right)

    • Cross-Modal Transformer: The difference-enhanced visual representations and textual tokens are concatenated and processed by a multi-layer self-attention Transformer FcrossF_{\text{cross}} (hidden dimension 512, 8 attention heads; 2 layers for Birds-to-Words, 3 layers for CLEVR-Change):

    V^(1),V^(2),T^=Fcross(V~(1),V~(2),T)\hat{V}^{(1)}, \hat{V}^{(2)}, \hat{T} = F_{\text{cross}}\left(\tilde{V}^{(1)}, \tilde{V}^{(2)}, T\right)

  2. Knowl 2 — Masked Visual Contrastive Learning Pre-training Objective

    model/method

    In the Masked Visual Contrastive Learning (MVCL) pre-training task, 15% of the input image grid features from one image in the pair {V(1),V(2)}\{V^{(1)}, V^{(2)}\} are masked by replacing them with zero vectors. Masking is restricted to only one image at a time so that the missing information can be recovered from the remaining image and the difference description TT.

    Instead of direct pixel or feature regression, MVCL optimizes a Noise Contrastive Estimation (NCE) loss:

    LMVCL=EV,T∈D[−log⁡exp⁡(d(vm,vm+)/τ1)exp⁡(d(vm,vm+)/τ1)+∑v′∈N(vm)exp⁡(d(vm,v′)/τ1)]\mathcal{L}_{\text{MVCL}} = \mathbb{E}_{V, T \in \mathcal{D}} \left[ -\log \frac{\exp\left(d(v_m, v_m^+)/\tau_1\right)}{\exp\left(d(v_m, v_m^+)/\tau_1\right) + \sum_{v' \in \mathcal{N}(v_m)} \exp\left(d(v_m, v')/\tau_1\right)} \right]

    where:

    • D\mathcal{D} is the training dataset and V={V(1),V(2)}V = \{V^{(1)}, V^{(2)}\}.
    • vmv_m is the cross-modal transformer output representation at the position of the masked visual token.
    • vm+v_m^+ is the original unmasked image feature at that position (positive target).
    • N(vm)\mathcal{N}(v_m) is the negative sample set, defined as all unmasked visual features in the mini-batch.
    • d(⋅,⋅)d(\cdot, \cdot) denotes cosine similarity.
    • τ1\tau_1 is a temperature hyperparameter set to 1.0.
  3. Knowl 3 — Fine-Grained Difference Aligning Pre-training Task

    model/method

    The Fine-grained Difference Aligning (FDA) task aligns cross-modal representations by contrasting a positive image-pair difference description (V,T+)(V, T^+) against a set of constructed hard negative sentences NT\mathcal{N}_T.

    The global visual representation is computed by mean-pooling the output representations of [IMG1][\text{IMG1}] and [IMG2][\text{IMG2}], while the textual representation is taken from the output of [CLS][\text{CLS}]. For each sample, 6 hard negative descriptions are constructed using three strategies (2 per strategy):

    1. Retrieve: The most similar difference descriptions T−T^- from other training samples are retrieved using TF-IDF similarity.
    2. Replace: Adjectives and nouns in T+T^+ are identified using Stanford CoreNLP, ranked by TF-IDF scores, and the top 50% most important words are replaced with randomly chosen words possessing identical POS tags.
    3. Confuse: Semantic relations are inverted by swapping subjects across sentences or switching the subject and object inside the sentence (e.g., transforming "animal1 is larger than animal2" into "animal2 is larger than animal1").

    The FDA loss objective is formulated as:

    LFDA=EV,T∈D[−log⁡exp⁡(d(V,T+)/τ2)exp⁡(d(V,T+)/τ2)+∑T−∈NTexp⁡(d(V,T−)/τ2)]\mathcal{L}_{\text{FDA}} = \mathbb{E}_{V, T \in \mathcal{D}} \left[ -\log \frac{\exp\left(d(V, T^+)/\tau_2\right)}{\exp\left(d(V, T^+)/\tau_2\right) + \sum_{T^- \in \mathcal{N}_T} \exp\left(d(V, T^-)/\tau_2\right)} \right]

    where d(⋅,⋅)d(\cdot, \cdot) is cosine similarity and τ2=1.0\tau_2 = 1.0 is the temperature hyperparameter.

  4. Knowl 4 — Causal Finetuning and Autoregressive Generation for Image Difference Captioning

    model/method

    To adapt the pre-trained bi-directional cross-modal transformer for difference description generation, the attention mechanism and training objective are modified during fine-tuning:

    • Finetuning Masking Schema: The model is fine-tuned using the Masked Language Modeling (MLM) objective. Text tokens are constrained with a uni-directional (causal) attention mask, restricting each token wtw_t to attend only to previous tokens w<tw_{<t}. Visual tokens retain full bi-directional attention, and text tokens can attend to all visual features from both images.

    • Autoregressive Inference: During inference, visual features {V(1),V(2)}\{V^{(1)}, V^{(2)}\} and the [CLS][\text{CLS}] token are fed together with the start token [BOS][\text{BOS}] and a [MASK][\text{MASK}] token. The model samples word w0w_0 via greedy decoding from the output distribution at the [MASK][\text{MASK}] position. At step tt, the prefix sequence {[BOS],w0,…,wt−1,[MASK]}\{[\text{BOS}], w_0, \dots, w_{t-1}, [\text{MASK}]\} is fed into the network to predict word wtw_t. Decoding repeats iteratively until the end token [EOS][\text{EOS}] is generated.

  5. Knowl 5 — Cross-Task Data Expansion Strategy for IDC

    model/method

    To address the limited availability of annotated triplet image difference data, the framework incorporates supervision from General Image Captioning (GIC) and Fine-Grained Visual Classification (FGVC):

    1. GIC Data Integration: GIC image-text pairs (image,text)(\text{image}, \text{text}) are adapted into triplets by pairing the real image with an empty image padded with zero vectors. In the cross-modal transformer, the padded zero vectors are masked out from self-attention. In the MVCL pre-training task, masking is restricted exclusively to tokens from the real image.

    2. FGVC Data Integration: Pairs (img1,img2)(img_1, img_2) are sampled such that 50% belong to the same category and 50% belong to different categories.

      • The single-image encoder FsingF_{\text{sing}} is updated on individual images using a fine-grained classification loss and a contrastive loss.
      • The pair-image encoder FpairF_{\text{pair}} receives representations of both images and is trained with a binary matching loss predicting whether img1img_1 and img2img_2 share the same class label, improving inter-image difference discrimination.
  6. Knowl 6 — Distractor Discrimination Objective for Change Captioning

    model/method

    In change captioning datasets such as CLEVR-Change, approximately 50% of image pairs represent distractor scenarios containing only viewpoint, camera zoom, or illumination shifts without actual semantic changes.

    To prevent the captioner from hallucinating changes on distractors, a distractor judging auxiliary task is jointly trained during fine-tuning. The final layer representations of [IMG1][\text{IMG1}] and [IMG2][\text{IMG2}] are concatenated and passed to a binary classifier that predicts whether the visual difference between the two images is an irrelevant distractor or a valid semantic change.

  7. Knowl 7 — Empirical Performance on CLEVR-Change Benchmark and Change Type Breakdown

    data/table

    The proposed framework was evaluated on the CLEVR-Change test set (7,970 image pairs) against existing change captioning approaches. Metrics include BLEU-4 (B4), METEOR (M), ROUGE-L (R), and CIDEr (C, the primary metric on CLEVR-Change).

    Model B4 M R C
    Capt-Dual-Att (2019) 43.5 32.7 - 108.5
    DUDA (2019) 47.3 33.9 - 112.0
    VAM (2020) 50.3 37.0 69.7 114.9
    VAM+ (2020) 51.3 37.8 70.4 115.8
    IFDC (2021a) 49.2 32.5 69.1 118.7
    DUDA+Aux (2021) 51.2 37.7 70.5 115.4
    Ours 51.2 36.2 71.7 128.9

    The breakdown of CIDEr performance across specific change categories on CLEVR-Change is as follows:

    Model Color Texture Move Add Drop Distractor
    DUDA 120.4 86.7 56.4 108.2 103.4 110.8
    VAM+ 122.1 98.7 82.0 126.3 115.8 122.6
    IFDC 133.2 99.1 82.1 128.2 118.5 114.2
    Ours 131.2 101.1 81.7 133.3 116.5 145.0

    The model improves overall CIDEr from the previous state-of-the-art of 118.7 to 128.9, with the largest gains occurring on Texture (101.1 vs 99.1), Add (133.3 vs 128.2), and Distractor pairs (145.0 vs 122.6).

  8. Knowl 8 — Empirical Performance on Birds-to-Words Benchmark with Cross-Task Data

    data/table

    The model was evaluated on the Birds-to-Words test set with and without cross-task data expansion using CUB (single-image bird captions) and NABirds (fine-grained bird classification). The primary evaluation metric for Birds-to-Words is ROUGE-L (R), alongside BLEU-4 (B4), METEOR (M), and CIDEr-D (C(D)).

    Model B4 M C(D) R
    Neural Naturalist (2019) 22.0 - 25.0 43.0
    Relational Speaker (2019) 21.5 22.4 5.8 43.4
    DUDA (2019) 23.9 21.9 4.6 44.3
    L2C (2021) 31.3 - 15.1 45.3
    L2C (+CUB) (2021) 31.8 - 16.3 45.6
    Ours (Birds-to-Words only) 28.0 23.1 18.6 48.4
    Ours (+Extra Data: CUB + NABirds) 31.0 23.4 25.3 49.1

    Ablation over cross-task data source combinations on Birds-to-Words (B2W):

    Data Configuration B4 M C(D) R
    B2W only 28.0 23.1 18.6 48.4
    B2W + CUB 29.3 23.1 23.8 48.5
    B2W + NABirds 27.5 23.3 21.9 48.5
    B2W + CUB + NABirds 31.0 23.4 25.3 49.1

    Even without extra data, the proposed method achieves 48.4 ROUGE-L, outperforming previous methods that utilized external CUB data (45.6). Adding both CUB and NABirds provides complementary background knowledge, boosting CIDEr-D to 25.3 and ROUGE-L to 49.1.

  9. Knowl 9 — Ablation of Pre-training Objectives and Architecture Modules on CLEVR-Change

    data/table

    An ablation study on the CLEVR-Change dataset evaluates the individual contributions of pre-training objectives (Masked Language Modeling [MLM], Masked Visual Contrastive Learning [MVCL], Fine-grained Difference Aligning [FDA]), the Image Difference Encoder (DE), and the fine-tuning Distractor Judging task.

    Configuration DE B4 M R C
    1. None (Finetuning only) ✓\checkmark 32.7 27.7 57.2 89.8
    2. MLM ✓\checkmark 36.7 28.2 60.9 94.9
    3. MLM + MVCL ✓\checkmark 50.3 37.6 70.6 119.7
    4. MLM + MVCL + FDA (Full model) ✓\checkmark 51.2 36.2 71.7 128.9
    5. MLM + MVCL + FDA w/o DE ×\times 49.2 35.8 68.8 107.9
    6. Full model w/o Distractor Judging ✓\checkmark 49.8 36.9 69.2 123.5

    Key findings include:

    • Adding visual contrastive pre-training (MVCL) produces the largest single performance leap, raising CIDEr from 94.9 to 119.7.
    • Adding hard negative cross-modal alignment (FDA) further raises CIDEr from 119.7 to 128.9.
    • Removing the Image Difference Encoder (DE) causes a 21.0 point drop in CIDEr (128.9 to 107.9).
    • Omitting the auxiliary distractor judging task during fine-tuning reduces CIDEr from 128.9 to 123.5.
  10. Knowl 10 — Repetition and Semantic Conflict in Generated Captions

    limitation

    Qualitative analysis of difference descriptions generated by the model indicates two recurring error patterns:

    1. Repetition: The model occasionally generates redundant clauses that describe the same visual difference repeatedly within a single output caption.
    2. Semantic Conflicts: The model sometimes overlooks conflicts between generated attributes or produces contradictory comparative statements across sentences for the two images.

Coverage note — No substantial contributed material was omitted. All pre-training objectives, architectural details, cross-task expansions, benchmark results, ablations, and error limitations are covered.

References

  1. 1.Chen, Y.-C.; Li, L.; Yu, L.; El Kholy, A.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020. Uniter: Universal image-text representation learning. In European conference on computer vision, 104–120. Springer.
  2. 2.Denkowski, M.; and Lavie, A. 2014. Meteor universal: Language specific translation evaluation for any target language. In Proceedings of the ninth workshop on statistical machine translation, 376–380.
  3. 3.Dubey, A.; Gupta, O.; Guo, P.; Raskar, R.; Farrell, R.; and Naik, N. 2018. Pairwise confusion for fine-grained visual classification. In Proceedings of the European conference on computer vision (ECCV), 70–86.
  4. 4.Fei, Z. 2021. Partially non-autoregressive image captioning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, 1309–1316.
  5. 5.Forbes, M.; Kaeser-Chen, C.; Sharma, P.; and Belongie, S. 2019. Neural Naturalist: Generating Fine-Grained Image Comparisons. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 708–717.
  6. 6.Ge, W.; Lin, X.; and Yu, Y. 2019. Weakly supervised complementary parts models for fine-grained image classification from the bottom up. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3034–3043.
  7. 7.Gu, J.; Cai, J.; Wang, G.; and Chen, T. 2018. Stack-captioning: Coarse-to-fine learning for image captioning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32.
  8. 8.He, J.; Chen, J.; Liu, S.; Kortylewski, A.; Yang, C.; Bai, Y.; Wang, C.; and Yuille, A. 2021. TransFG: A Transformer Architecture for Fine-grained Recognition. arXiv preprint arXiv:2103.07976.
  9. 9.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 770–778.
  10. 10.Hosseinzadeh, M.; and Wang, Y. 2021. Image Change Captioning by Learning From an Auxiliary Task. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2725–2734.
  11. 11.Hu, X.; Yin, X.; Lin, K.; Zhang, L.; Gao, J.; Wang, L.; and Liu, Z. 2021. VIVO: Visual Vocabulary Pre-Training for Novel Object Captioning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, 1575–1583.
  12. 12.Huang, Q.; Liang, Y.; Wei, J.; Yi, C.; Liang, H.; Leung, H.-f.; and Li, Q. 2021a. Image Difference Captioning with Instance-Level Fine-Grained Feature Representation. IEEE Transactions on Multimedia, 1–1.
  13. 13.Huang, Z.; Zeng, Z.; Huang, Y.; Liu, B.; Fu, D.; and Fu, J. 2021b. Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12976–12985.
  14. 14.Huo, Y.; Zhang, M.; Liu, G.; Lu, H.; Gao, Y.; Yang, G.; Wen, J.; Zhang, H.; Xu, B.; Zheng, W.; et al. 2021. WenLan: Bridging vision and language by large-scale multi-modal pre-training. arXiv preprint arXiv:2103.06561.
  15. 15.Jhamtani, H.; and Berg-Kirkpatrick, T. 2018. Learning to Describe Differences Between Pairs of Similar Images. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 4024–4034.
  16. 16.Jiang, W.; Ma, L.; Chen, X.; Zhang, H.; and Liu, W. 2018. Learning to guide decoding for image captioning. In Thirty-second AAAI conference on artificial intelligence.
  17. 17.Kim, W.; Son, B.; and Kim, I. 2021. ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision. In ICML.
  18. 18.Kingma, D. P.; and Ba, J. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  19. 19.Lee, H.; Yoon, S.; Dernoncourt, F.; Bui, T.; and Jung, K. 2021. UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning. arXiv preprint arXiv:2106.14019.
  20. 20.Li, G.; Duan, N.; Fang, Y.; Gong, M.; and Jiang, D. 2020a. Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 11336–11344.
  21. 21.Li, L.; Chen, Y.-C.; Cheng, Y.; Gan, Z.; Yu, L.; and Liu, J. 2020b. Hero: Hierarchical Encoder for Video+ Language Omni-representation Pre-training. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2046–2065.
  22. 22.Li, W.; Gao, C.; Niu, G.; Xiao, X.; Liu, H.; Liu, J.; Wu, H.; and Wang, H. 2020c. UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning. arXiv preprint arXiv:2012.15409.
  23. 23.Li, X.; Yin, X.; Li, C.; Zhang, P.; Hu, X.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; et al. 2020d. Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision, 121–137. Springer.
  24. 24.Lin, C.-Y. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, 74–81.
  25. 25.Liu, C.; Xie, H.; Zha, Z.-J.; Ma, L.; Yu, L.; and Zhang, Y. 2020. Filtration and distillation: Enhancing region attention for fine-grained visual categorization. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 11555–11562.
  26. 26.Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019. ViLBERT: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, 13–23.
  27. 27.Luo, H.; Ji, L.; Shi, B.; Huang, H.; Duan, N.; Li, T.; Li, J.; Bharti, T.; and Zhou, M. 2020. UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation. arXiv preprint arXiv:2002.06353.
  28. 28.Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 311–318.
  29. 29.Park, D. H.; Darrell, T.; and Rohrbach, A. 2019. Robust change captioning. In Proceedings of the IEEE international conference on computer vision, 4624–4633.
  30. 30.Qi, D.; Su, L.; Song, J.; Cui, E.; Bharti, T.; and Sacheti, A. 2020. Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data. arXiv preprint arXiv:2001.07966.
  31. 31.Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020.
  32. 32.Rennie, S. J.; Marcheret, E.; Mroueh, Y.; Ross, J.; and Goel, V. 2017. Self-critical sequence training for image captioning. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7008–7024.
  33. 33.Shi, X.; Yang, X.; Gu, J.; Joty, S.; and Cai, J. 2020. Finding it at another side: A viewpoint-adapted matching encoder for change captioning. In European Conference on Computer Vision, 574–590. Springer.
  34. 34.Sun, C.; Baradel, F.; Murphy, K.; and Schmid, C. 2019. Learning video representations using contrastive bidirectional transformer. arXiv preprint arXiv:1906.05743.
  35. 35.Tan, H.; Dernoncourt, F.; Lin, Z.; Bui, T.; and Bansal, M. 2019. Expressing Visual Relationships via Language. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 1873–1883.
  36. 36.Van Horn, G.; Branson, S.; Farrell, R.; Haber, S.; Barry, J.; Ipeirotis, P.; Perona, P.; and Belongie, S. J. 2015. Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection. In CVPR.
  37. 37.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. In Advances in neural information processing systems, 5998–6008.
  38. 38.Vedantam, R.; Lawrence Zitnick, C.; and Parikh, D. 2015. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 4566–4575.
  39. 39.Vinyals, O.; Toshev, A.; Bengio, S.; and Erhan, D. 2015. Show and tell: A neural image caption generator. In Proceedings of the IEEE conference on computer vision and pattern recognition, 3156–3164.
  40. 40.Wah, C.; Branson, S.; Welinder, P.; Perona, P.; and Belongie, S. 2011. The Caltech-UCSD Birds-200-2011 Dataset. Technical Report CNS-TR-2011-001, California Institute of Technology.
  41. 41.Xu, K.; Ba, J.; Kiros, R.; Cho, K.; Courville, A.; Salakhudinov, R.; Zemel, R.; and Bengio, Y. 2015. Show, attend and tell: Neural image caption generation with visual attention. In International conference on machine learning, 2048–2057. PMLR.
  42. 42.Yan, A.; Wang, X.; Fu, T.-J.; and Wang, W. Y. 2021. L2C: Describing Visual Differences Needs Semantic Understanding of Individuals. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, 2315–2320.
  43. 43.Zhao, W.; Wu, X.; and Zhang, X. 2020. Memcap: Memorizing style knowledge for image captioning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 12984–12992.
  44. 44.Zhou, L.; Palangi, H.; Zhang, L.; Hu, H.; Corso, J.; and Gao, J. 2020. Unified vision-language pre-training for image captioning and vqa. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 13041–13049.

Citation

MLA
Yao, L., et al. “Image Difference Captioning with Pre-training and Contrastive Learning”. arXiv, 2022, http://arxiv.org/abs/2202.04298v1.
APA
Yao, L., Wang, W., & Jin, Q. (2022). Image Difference Captioning with Pre-training and Contrastive Learning. arXiv. http://arxiv.org/abs/2202.04298v1
Chicago
Yao, L., W. Wang, and Q. Jin. 2022. “Image Difference Captioning with Pre-training and Contrastive Learning”. arXiv. http://arxiv.org/abs/2202.04298v1.
Harvard
Yao, L., Wang, W. and Jin, Q. (2022) “Image Difference Captioning with Pre-training and Contrastive Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2202.04298v1.
Vancouver
1. Yao L, Wang W, Jin Q (2022) Image Difference Captioning with Pre-training and Contrastive Learning. arXiv

BibTeX

@article{yao2022image,
  title = {Image Difference Captioning with Pre-training and Contrastive Learning},
  author = {Yao, Linli and Wang, Weiying and Jin, Qin},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2202.04298v1},
  eprint = {2202.04298}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF