GeneCIS: A Benchmark for General Conditional Image Similarity

Sagar VazeNicolas CarionIshan Misra

article2023CVPR67 citations

Introduces the GeneCIS benchmark to evaluate zero-shot conditional image similarity and presents a scalable training strategy that mines image-caption datasets to significantly improve retrieval models when adapting to open-ended similarity criteria.

Listen

Standard computer vision models typically rely on a single, fixed notion of image similarity, such as matching overall object categories. However, practical human workflows require dynamic similarity comparisons based on context, such as identifying items with the same color, isolating a specific object in a complex scene, or modifying a single attribute. The article addresses this limitation by formalizing the task of general conditional image similarity, where models must retrieve relevant images based on an explicit user-specified text prompt without task-specific retraining.

The main objective of the article is to establish a rigorous evaluation benchmark for conditional image similarity and introduce a scalable, automated training approach that enables models to adapt to diverse visual conditions in an open-set, zero-shot setting.

To accomplish this, the authors introduced the GeneCIS benchmark, comprising four distinct evaluation tasks constructed from existing visual datasets: focusing on an attribute, changing an attribute, focusing on an object, and changing an object. Each task evaluates a model's ability to select the correct target image from small galleries containing challenging distractor images. To train models without labor-intensive manual labeling, the authors developed an automated pipeline that mines 1.6 million training triplets from 3 million web image-caption pairs. This method parses captions into subject-predicate-object relationships, filters them for visual concreteness using an external linguistic database, and pairs images sharing common subjects under modified conditions using a contrastive learning architecture.

The investigation produced several key findings. First, established vision-language models struggle significantly on conditional retrieval; simple image-and-text averaging achieved only a 12.6% average top-1 recall across the GeneCIS tasks. Second, performance on conditional similarity is only weakly correlated with standard benchmark accuracy: a 10% gain in zero-shot ImageNet accuracy yielded approximately a 1% gain on GeneCIS. Third, the proposed automated caption-mining approach achieved a 16.8% average top-1 recall, outperforming all zero-shot baselines as well as models trained on manually curated datasets. Finally, when tested on external composed image retrieval benchmarks, the zero-shot model surpassed supervised state-of-the-art models on the MIT-States benchmark (15.8% vs. 15.6% top-1 recall) and outperformed zero-shot baselines on the CIRR benchmark (27.3% vs. 21.8% top-1 recall).

These findings demonstrate that general vision systems cannot rely solely on standard scaling to solve complex, instruction-based retrieval tasks. Instead, explicit conditioning mechanisms are required. By leveraging existing, abundant image-caption data, organizations can significantly enhance multi-modal search and interactive computer vision applications without the high costs and timelines associated with bespoke manual data annotation.

Organizations developing fine-grained visual search or interactive media systems should adopt automated relationship parsing on large caption datasets rather than investing in manual condition labeling. Practitioners should also evaluate their vision backbones directly on instruction-based benchmarks rather than relying on general classification metrics to predict conditional retrieval performance. Further work should explore scaling this automated mining technique to web-scale datasets containing billions of image-text pairs.

Confidence in these findings is supported by consistent gains across multiple distinct benchmarks and ablation studies. However, the initial GeneCIS release (version 0) contains minor underlying label noise inherited from source datasets, and the current training triplet distribution exhibits a natural bias toward object-modification tasks rather than attribute adjustments. Users should account for these boundary conditions when deploying the approach across specialized domains.

arXiv: 2306.07969
Cover for GeneCIS: A Benchmark for General Conditional Image Similarity

Abstract

We argue that there are many notions of ‘similarity’ and that models, like humans, should be able to adapt to these dynamically. This contrasts with most representation learning methods, supervised or self-supervised, which learn a fixed embedding function and hence implicitly assume a single notion of similarity. For instance, models trained on ImageNet are biased towards object categories, while a user might prefer the model to focus on colors, textures or specific elements in the scene. In this paper, we propose the GeneCIS (‘genesis’) benchmark, which measures models’ ability to adapt to a range of similarity conditions. Extending prior work, our benchmark is designed for zero-shot evaluation only, and hence considers an open-set of similarity conditions. We find that baselines from powerful CLIP models struggle on GeneCIS and that performance on the benchmark is only weakly correlated with ImageNet accuracy, suggesting that simply scaling existing methods is not fruitful. We further propose a simple, scalable solution based on automatically mining information from existing image-caption datasets. We find our method offers a substantial boost over the baselines on GeneCIS, and further improves zero-shot performance on related image retrieval benchmarks. In fact, though evaluated zero-shot, our model surpasses state-of-the-art supervised models on MIT-States.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Conditional Similarity
  • 3.1. Challenges in training and evaluation
  • 4. The GeneCIS Benchmark
  • 5. Method
  • 5.1. Preliminaries
  • 5.2. Scalable training for conditional similarity
  • 6. Main Experiments
  • 6.1. Baselines and Specific Solutions for GeneCIS
  • 6.2. Implementation Details
  • 6.3. Analysis on GeneCIS
  • 6.4. Comparisons to Prior Work
  • 7. Analysis
  • 8. Conclusion
  • References

Knowls

  1. Knowl 1 — GeneCIS Benchmark for General Conditional Image Similarity

    experimental setup

    The GeneCIS (General Conditional Image Similarity) benchmark evaluates vision-language models on their ability to adapt dynamically to diverse, user-specified similarity conditions in a zero-shot retrieval setting.

    GeneCIS spans two primary dimensions of similarity: condition scope (focusing on an existing element vs. specifying a negative change/modification) and semantic entity (attribute vs. scene object). This yields four distinct retrieval tasks:

    1. Focus on an Attribute: Evaluates isolating a specific attribute dimension (e.g., condition c="color"c = \text{"color"}). Given a reference image IRI^R showing an object (e.g., a white laptop), the model must retrieve the gallery image containing an object with the identical attribute value. Sourced from the Visual Attributes in the Wild (VAW) dataset (2,000 templates, gallery size 10).
    2. Change an Attribute: Evaluates negative similarity conditioning. Given a reference image IRI^R and a replacement attribute condition cc (e.g., c="olive green"c = \text{"olive green"} for a green train), the target ITI^T is the same object category displaying the specified modified attribute. Sourced from VAW (2,112 templates, gallery size 15).
    3. Focus on an Object: Given a complex reference scene IRI^R with multiple objects and a condition cc highlighting a single object (e.g., c="refrigerator"c = \text{"refrigerator"}), the positive target ITI^T contains the condition object embedded in a matching scene context. Sourced from COCO Panoptic Segmentation (1,960 templates, gallery size 15).
    4. Change an Object: Evaluates condition-based scene modification. Given a reference scene IRI^R and a condition specifying an object to add (e.g., c="ceiling"c = \text{"ceiling"}), the positive target ITI^T depicts the same scene with the requested object present. Sourced from COCO Panoptic Segmentation (1,960 templates, gallery size 15).

    Each retrieval gallery contains exactly one positive target image and 99 to 1414 carefully curated distractor images that share implicit similarity with the reference (e.g., matching scene but missing the condition object; matching object but wrong attribute) or the condition alone (e.g., condition object in an unrelated scene), requiring models to jointly utilize both the visual reference and the textual condition.

  2. Knowl 2 — Conditional Image Similarity Architecture and Contrastive Optimization

    model/method

    Conditional image similarity is formulated as computing a scalar similarity f(IT;IR,c)f(I^T; I^R, c) between a target image ITI^T and a reference image IRI^R subject to an explicit conditioning text prompt cc.

    The framework encodes image and text inputs via pretrained encoders:

    • Reference image embedding: xR=Φ(IR)∈RDx^R = \Phi(I^R) \in \mathbb{R}^D
    • Target image embedding: xT=Φ(IT)∈RDx^T = \Phi(I^T) \in \mathbb{R}^D
    • Condition text embedding: e=Ψ(c)∈RDe = \Psi(c) \in \mathbb{R}^D

    where Φ(⋅)\Phi(\cdot) and Ψ(⋅)\Psi(\cdot) denote visual and textual backbones initialized from CLIP. A non-linear Combiner network g(xR,e):RD×RD→RDg(x^R, e): \mathbb{R}^D \times \mathbb{R}^D \to \mathbb{R}^D fuses the reference visual feature and condition text feature into a conditioned query vector. The scalar similarity score is defined as the inner product:

    f(IT;IR,c)=g(xR,e)⋅xTf(I^T; I^R, c) = g(x^R, e) \cdot x^T

    Given a mini-batch of triplets B={(IiR,IiT,ci)}i=1∣B∣B = \{(I_i^R, I_i^T, c_i)\}_{i=1}^{|B|}, the parameters of the encoders and Combiner (Φ,Ψ,g)(\Phi, \Psi, g) are optimized end-to-end using a symmetrical InfoNCE contrastive loss with temperature τ\tau:

    L=−1∣B∣∑i∈Blog⁡exp⁡(g(xiR,ei)⋅xiT/τ)∑j∈Bexp⁡(g(xiR,ei)⋅xjT/τ)\mathcal{L} = -\frac{1}{|B|} \sum_{i \in B} \log \frac{\exp\left(g(x_i^R, e_i) \cdot x_i^T / \tau\right)}{\sum_{j \in B} \exp\left(g(x_i^R, e_i) \cdot x_j^T / \tau\right)}

  3. Knowl 3 — Automatic Mining of Conditional Similarity Triplets from Image Captions

    algorithm

    To train open-set conditional similarity models without manual annotation, training triplets (IR,IT,c)(I^R, I^T, c) are extracted automatically from large-scale image-caption datasets by parsing semantic scene graphs and filtering by visual concreteness.

    Input: Image-caption dataset D = {(I_k, text_k)}, Concreteness threshold theta = 4.8
    Output: Mined triplet dataset D_train = {(I^R, I^T, c)}
    for each (I_k, text_k) in D do
        relationships[k] = ParseTextToSceneGraph(text_k) // Extracts (Subject, Predicate, Object) tuples
        filtered_rel[k] = []
        for each (S, P, O) in relationships[k] do
            score_S = LookupConcreteness(S)
            score_O = LookupConcreteness(O)
            if (score_S + score_O) / 2 >= theta then
                filtered_rel[k].append((S, P, O))
            end if
        end for
    end for
    D_train = []
    for each reference image I^R with relation (S_R, P_R, O_R) in filtered_rel do
        candidates = all relations (I^T, (S_T, P_T, O_T)) where S_T == S_R and O_T != O_R
        if candidates is not empty then
            Sample target relation (I^T, (S_T, P_T, O_T)) from candidates
            c = Concatenate(P_T, " ", O_T)
            D_train.append((I^R, I^T, c))
        end if
    end for
    return D_train

    Concreteness scores are looked up from the Brysbaert et al. English lemma database (ratings from 1 to 5). Filtering prevents abstract non-visual entities (e.g., pronouns or time terms) from forming invalid conditions.

  4. Knowl 4 — Evaluation Results on the GeneCIS Benchmark

    data/table

    The table below compares the performance of baseline models, task-specific open-vocabulary solutions, and the trained Combiner model on the GeneCIS benchmark across all four evaluation tasks. Performance is measured using Recall@K (R@K, in %) for K∈{1,2,3}K \in \{1, 2, 3\} and Average Recall@1 across tasks.

    Method Focus Attribute Change Attribute Focus Object Change Object Average
    R@1 R@2 R@3 R@1 R@2 R@3 R@1 R@2 R@3 R@1 R@2 R@3 R@1
    Specific Solution (Focus Attr.) 20.8 32.6 41.1 - - - - - - - - - -
    Specific Solution (Change Attr.) - - - 15.2 25.8 35.6 - - - - - - -
    Specific Solution (Object) - - - - - - 18.7 30.3 37.4 18.1 28.7 34.5 -
    Image Only 17.7 30.9 41.9 11.9 20.8 28.8 9.3 18.2 26.2 7.2 16.7 24.9 11.5
    Text Only 10.2 20.5 29.6 9.5 17.6 26.4 6.5 16.8 22.4 6.2 13.9 21.4 8.1
    Image + Text 15.6 26.3 37.1 12.6 22.9 32.0 10.8 21.0 31.2 11.3 21.5 30.3 12.6
    Combiner (CIRR) 15.1 27.7 39.8 12.1 22.8 31.8 13.5 25.4 36.7 15.4 28.0 39.6 14.0
    Combiner (CC3M, Ours) 19.0 31.0 41.5 16.6 27.5 36.5 14.7 25.9 36.1 16.8 29.1 39.7 16.8

    The Combiner trained on 1.6M triplets mined from Conceptual Captions 3M (CC3M) with a ResNet-50x4 CLIP backbone achieves an Average R@1 of 16.8%, outperforming the Image Only (11.5%), Text Only (8.1%), feature averaging baseline Image + Text (12.6%), and the model trained on human-annotated CIRR data (14.0%). The unimodal baselines struggle because distractor images require joint conditioning on both the reference image and the text condition.

  5. Knowl 5 — Zero-Shot Performance on MIT-States and CIRR Benchmarks

    empirical result

    The Combiner model trained on automatically mined CC3M triplets generalizes zero-shot to established Composed Image Retrieval (CIR) benchmarks (MIT-States and CIRR), achieving performance competitive with or superior to supervised models trained directly on those benchmarks.

    On the MIT-States test set (global retrieval over the full dataset):

    • Zero-shot Combiner (CC3M): Recall@1 = 15.8%, Recall@5 = 37.5%, Recall@10 = 49.4%
    • Zero-shot Image + Text baseline: Recall@1 = 13.3%, Recall@5 = 31.7%, Recall@10 = 42.6%
    • Zero-shot Text Only baseline: Recall@1 = 9.5%, Recall@5 = 22.5%, Recall@10 = 31.4%
    • Zero-shot Image Only baseline: Recall@1 = 3.7%, Recall@5 = 14.1%, Recall@10 = 22.9%
    • Supervised methods (trained on MIT-States): MAN (15.6% R@1), HCL (15.2% R@1), LBF (14.7% R@1), ComposeAE (13.9% R@1), TIRG (12.2% R@1).

    On the CIRR official test server (global retrieval over the full gallery):

    • Zero-shot Combiner (CC3M): Recall@1 = 27.3%, Recall@5 = 57.0%, Recall@10 = 71.1%
    • Zero-shot Image + Text baseline: Recall@1 = 21.8%, Recall@5 = 50.9%, Recall@10 = 63.7%
    • Zero-shot Text Only baseline: Recall@1 = 20.7%, Recall@5 = 43.9%, Recall@10 = 56.1%
    • Zero-shot Image Only baseline: Recall@1 = 7.5%, Recall@5 = 23.9%, Recall@10 = 34.7%
    • Supervised methods (trained on CIRR): Combiner (CIRR, standard) achieves 38.5% R@1, Combiner (CIRR, fine-tuned backbones) achieves 40.9% R@1, while CIRPLANT (19.6% R@1) and ARTEMIS (17.0% R@1) underperform the zero-shot Combiner (CC3M).
  6. Knowl 6 — Ablation Analysis of Concreteness Filtering and Backbone Fine-Tuning

    data/table

    The ablation study below evaluates the impact of training pipeline components on the GeneCIS benchmark, measured by Average Recall@1 (in %):

    Model Configuration Average Recall @ 1 (%)
    Full Model (CC3M, Concreteness ≥4.8\ge 4.8, Fine-tuned Backbone) 16.8
    No filtering for visual concreteness 15.0
    Freezing CLIP image backbone 14.7
    Freezing CLIP text backbone 15.8
    Freezing entire CLIP backbone 15.1
    Training on SBU Captions instead of CC3M 16.5

    Removing the visual concreteness filter causes a 1.8% drop in Average R@1 (from 16.8% to 15.0%), demonstrating the necessity of discarding abstract grammatical relationships. Freezing the vision and text backbones reduces performance by 2.1% and 1.0% respectively, confirming that end-to-end adaptation of visual and linguistic features is critical. Changing the pretraining caption corpus from CC3M (3M captions) to SBU Captions (1M captions) yields a minor drop to 16.5%, confirming that the triplet mining pipeline is robust across different caption sources.

  7. Knowl 7 — Weak Correlation Between ImageNet Zero-Shot Accuracy and GeneCIS Performance

    empirical result

    Across multiple CLIP visual backbones (ResNet-50, ResNet-101, ResNet-50x4, ResNet-50x16, ViT-B/32, ViT-B/16), model performance on the GeneCIS benchmark is weakly correlated with zero-shot ImageNet Top-1 classification accuracy.

    A 10% gain in zero-shot Top-1 accuracy on ImageNet translates to only an approximate 1% increase in Average Recall@1 on GeneCIS. Furthermore, training the Combiner model on mined conditional triplets provides a substantially larger gain over the unconditioned Image + Text baseline than any performance increase achievable simply by upgrading or scaling the underlying CLIP backbone architecture.

  8. Knowl 8 — Data Scaling Characteristics of Automatically Mined Triplets

    empirical result

    Model performance on the GeneCIS benchmark scales consistently with the volume of automatically mined training triplets. Evaluating across triplet counts scaled by successive factors of four (2.5×1042.5 \times 10^4, 1.0×1051.0 \times 10^5, 4.0×1054.0 \times 10^5, and 1.6×1061.6 \times 10^6 mined triplets from CC3M):

    • With concreteness filtering (extscore≥4.8 ext{score} \ge 4.8): GeneCIS Average Recall@1 monotonically increases from approximately 15.6% at 2.5×1042.5 \times 10^4 triplets to 16.8% at 1.6×1061.6 \times 10^6 triplets.
    • Without concreteness filtering: GeneCIS Average Recall@1 rises from ~14.0% at 2.5×1042.5 \times 10^4 to a peak of ~15.2% at 4.0×1054.0 \times 10^5, but saturates and degrades to 15.0% at 1.6×1061.6 \times 10^6 triplets.

    This demonstrates that automatic triplet mining scales effectively with dataset size, but visual concreteness filtering is necessary to prevent accumulation of noise from ungrounded abstract nouns at larger data scales.

Coverage note — No substantial contributed material was omitted. Implementation details of the task-specific open-vocabulary baselines (CLIP classification for attributes, Detic for objects) are summarized in the benchmark evaluation knowl.

References

  1. 1.Muhammad Umer Anwaar, Egor Labintcev, and Martin Kleinsteuber. Compositional learning of image-text query for image retrieval. In WACV, 2021. 2, 7
  2. 2.Omri Avrahami, Dani Lischinski, and Ohad Fried. Blended diffusion for text-driven editing of natural images. In CVPR, 2022. 4
  3. 3.Alberto Baldrati, Marco Bertini, Tiberio Uricchio, and Alberto Del Bimbo. Conditioned and composed image retrieval combining and partially fine-tuning clip-based features. In CVPRW, 2022. 2, 5, 6, 7, 8, 3
  4. 4.Maxim Berman, Hervé Jégou, Andrea Vedaldi, Iasonas Kokkinos, and Matthijs Douze. Multigrain: a unified image embedding for classes and instances. arXiv, 2019. 2
  5. 5.Tim Brooks, Aleksander Holynski, and Alexei A. Efros. Instructpix2pix: Learning to follow image editing instructions. ECCV, 2022. 4
  6. 6.Andrew Brown, Cheng-Yang Fu, Omkar Parkhi, Tamara L Berg, and Andrea Vedaldi. End-to-end visual editing with a generatively pre-trained artist. ECCV, 2022. 4
  7. 7.Andrew Brown, Weidi Xie, Vicky Kalogeiton, and Andrew Zisserman. Smooth-ap: Smoothing the path towards large-scale image retrieval. In ECCV, 2020. 2
  8. 8.Marc Brysbaert, Amy Beth Warriner, and Victor Kuperman. Concreteness ratings for 40 thousand generally known english word lemmas. In Behavior research methods, 2014. 6
  9. 9.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. NeurIPS, 2020. 1, 2
  10. 10.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, 2021. 1, 2
  11. 11.Jun-Cheng Chen, Vishal M. Patel, and Rama Chellappa. Unconstrained face verification using deep cnn features. In WACV, 2016. 1, 2
  12. 12.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020. 1, 2
  13. 13.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv, 2020. 1, 2
  14. 14.Katherine Crowson, Stella Biderman, Daniel Kornis, Dashiell Stander, Eric Hallahan, Louis Castricato, and Edward Raff. Vqgan-clip: Open domain image generation and editing with natural language guidance. In ECCV. Springer, 2022. 4
  15. 15.Yin Cui, Zeqi Gu, Dhruv Mahajan, Laurens Van Der Maaten, Serge Belongie, and Ser-Nam Lim. Measuring dataset granularity. arXiv, 2019. 2
  16. 16.Ginger Delmas, Rafael S Rezende, Gabriela Csurka, and Diane Larlus. Artemis: Attention-based retrieval with text-explicit matching and implicit similarity. In ICLR, 2022. 2, 7
  17. 17.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. 2, 8
  18. 18.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2020. 8, 3
  19. 19.Zhixiao Fu, Xinyuan Chen, Jianfeng Dong, and Shouling Ji. Multi-order adversarial representation learning for composed query image retrieval. In ICASSP, 2021. 7
  20. 20.Rinon Gal, Or Patashnik, Haggai Maron, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. Stylegan-nada: Clip-guided domain adaptation of image generators. ACM Transactions on Graphics (TOG), 2022. 4
  21. 21.Robert L. Goldstone and Ji Yun Son. 155 Similarity. Oxford University Press, 2012. 1
  22. 22.Agrim Gupta, Piotr Dollar, and Ross Girshick. LVIS: A dataset for large vocabulary instance segmentation. In CVPR, 2019. 4
  23. 23.Xintong Han, Zuxuan Wu, Phoenix X. Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S. Davis. Automatic spatially-aware fashion concept discovery. In ICCV, 2017. 2, 7
  24. 24.Xintong Han, Zuxuan Wu, Yu-Gang Jiang, and Larry S Davis. Learning fashion compatibility with bidirectional lstms. In ACM, 2017. 7
  25. 25.Bing He, Jia Li, Yifan Zhao, and Yonghong Tian. Part-regularized near-duplicate vehicle re-identification. In CVPR, June 2019. 2
  26. 26.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 8
  27. 27.Mehrdad Hosseinzadeh and Yang Wang. Composed query image retrieval using locally bounded features. In CVPR, 2020. 7
  28. 28.Phillip Isola, Joseph J. Lim, and Edward H. Adelson. Discovering states and transformations in image collections. In CVPR, 2015. 2, 7
  29. 29.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, 2021. 2
  30. 30.Mahmut Kaya and Hasan Ş akir Bilge. Deep metric learning: A survey. Symmetry, 11, 2019. 1, 2
  31. 31.Sultan Daud Khan and Habib Ullah. A survey of advances in vision-based vehicle re-identification. Computer Vision and Image Understanding, 2019. 2
  32. 32.Donghyun Kim, Kuniaki Saito, Samarth Mishra, Stan Sclaroff, Kate Saenko, and Bryan A. Plummer. Self-supervised visual attribute learning for fashion compatibility. ICCV Workshops, 2021. 2
  33. 33.Gwanghyun Kim, Taesung Kwon, and Jong Chul Ye. Diffusionclip: Text-guided diffusion models for robust image manipulation. In CVPR, 2022. 4
  34. 34.Sungyeon Kim, Dongwon Kim, Minsu Cho, and Suha Kwak. Proxy anchor loss for deep metric learning. In CVPR, 2020. 1, 2
  35. 35.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 3
  36. 36.Alex Kirillov, Tsung-Yi Lin, Holger Caesar, Ross Girshick, and Piotr Dollaŕ. Microsoft coco: Panoptic segmentation challenge, 2017. 4, 5, 7, 1, 2, 3
  37. 37.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations, 2016. 4, 5, 2
  38. 38.Gihyun Kwon and Jong Chul Ye. Clipstyler: Image style transfer with a single text condition. In CVPR, 2022. 4
  39. 39.Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and René Ranftl. Language-driven semantic segmentation. ICLR, 2022. 2, 8
  40. 40.Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollaŕ. Microsoft coco: Common objects in context. In ECCV, 2014. 4, 1, 2
  41. 41.Yen-Liang Lin, Son Tran, and Larry S. Davis. Fashion outfit complementary item retrieval. In CVPR, June 2020. 2
  42. 42.Xinchen Liu, Wu Liu, Huadong Ma, and Huiyuan Fu. Large-scale vehicle re-identification in urban surveillance videos. In ICME, 2016. 2
  43. 43.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021. 4
  44. 44.Zheyuan Liu, Cristian Rodriguez, Damien Teney, and Stephen Gould. Image retrieval on real-life images with pretrained vision-and-language models. In ICCV, 2021. 2, 6, 7, 3
  45. 45.Matthias Minderer, Alexey Gritsenko, Austin Stone, Maxim Neumann, Dirk Weissenborn, Alexey Dosovitskiy, Aravindh Mahendran, Anurag Arnab, Mostafa Dehghani, Zhuoran Shen, et al. Simple open-vocabulary object detection with vision transformers. ECCV, 2022. 2, 8
  46. 46.Samarth Mishra, Zhongping Zhang, Yuan Shen, Ranjitha Kumar, Venkatesh Saligrama, and Bryan A. Plummer. Effectively leveraging attributes for visual similarity. In ICCV, 2021. 1, 2, 3, 7
  47. 47.Ishan Misra, Abhinav Gupta, and Martial Hebert. From red wine to red tomato: Composition with context. In CVPR, 2017. 2
  48. 48.Andrei Neculai, Yanbei Chen, and Zeynep Akata. Probabilistic compositional embeddings for multimodal image retrieval. In CVPR, 2022. 4
  49. 49.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv, 2021. 4
  50. 50.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv, 2018. 5
  51. 51.Vicente Ordonez, Girish Kulkarni, and Tamara Berg. Im2text: Describing images using 1 million captioned photographs. In NeurIPS, 2011. 5, 8
  52. 52.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. NeurIPS, 2019. 3
  53. 53.Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. Styleclip: Text-driven manipulation of stylegan imagery. In ICCV, 2021. 4
  54. 54.Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. Film: Visual reasoning with a general conditioning layer. In AAAI, 2018. 2
  55. 55.Julia Peyre, Ivan Laptev, Cordelia Schmid, and Josef Sivic. Detecting unseen visual relations using analogies. In ICCV, 2019. 5
  56. 56.Khoi Pham, Kushal Kafle, Zhe Lin, Zhihong Ding, Scott Cohen, Quan Tran, and Abhinav Shrivastava. Learning to predict visual attributes in the wild. In CVPR, 2021. 2, 4, 5, 1, 3
  57. 57.Bryan A. Plummer, Liwei Wang, Christopher M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. IJCV, 2017. 4
  58. 58.Karl Popper. Conjectures and Refutations: The Growth of Scientific Knowledge. Routledge, 1963. 1
  59. 59.Filip Radenović, Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, and Ondŕej Chum. Revisiting oxford and paris: Large-scale image retrieval benchmarking. In CVPR, June 2018. 2
  60. 60.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In ICML, 2021. 2, 5, 6, 7, 3, 4
  61. 61.Jerome Revaud, Jon Almazan, Rafael S. Rezende, and Cesar Roberto de Souza. Learning with average precision: Training image retrieval with a listwise loss. In ICCV, October 2019. 2
  62. 62.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 4
  63. 63.Karsten Roth, Timo Milbich, Samarth Sinha, Prateek Gupta, Björn Ommer, and Joseph Paul Cohen. Revisiting training strategies and generalization performance in deep metric learning. In ICML, 2020. 1, 2
  64. 64.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa R Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. Laion-5b: An open large-scale dataset for training next generation image-text models. In NeurIPS Datasets and Benchmarks, 2022. 2, 8
  65. 65.Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D Manning. Generating semantically precise scene graphs from textual descriptions for improved image retrieval. In Proceedings of the fourth workshop on vision and language, 2015. 5
  66. 66.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL, 2018. 2, 5, 6, 8, 4, 11
  67. 67.Yi Sun, Yuheng Chen, Xiaogang Wang, and Xiaoou Tang. Deep learning face representation by joint identification-verification. In NeurIPS, 2014. 2
  68. 68.Yaniv Taigman, Ming Yang, Marc’Aurelio Ranzato, and Lior Wolf. Deepface: Closing the gap to human-level performance in face verification. In CVPR, June 2014. 1, 2
  69. 69.Reuben Tan, Mariya I. Vasileva, Kate Saenko, and Bryan A. Plummer. Learning similarity conditions without explicit supervision. In ICCV, 2019. 1, 2
  70. 70.Hugo Touvron, Alexandre Sablayrolles, Matthijs Douze, Matthieu Cord, and Hervé Jégou. Grafit: Learning fine-grained image representations with coarse labels. In ICCV, 2021. 2
  71. 71.Mariya I Vasileva, Bryan A Plummer, Krishna Dusad, Shreya Rajpal, Ranjitha Kumar, and David Forsyth. Learning type-aware embeddings for fashion compatibility. In ECCV, 2018. 7
  72. 72.Vijay Vasudevan, Benjamin Caine, Raphael Gontijo-Lopes, Sara Fridovich-Keil, and Rebecca Roelofs. When does dough become a bagel? analyzing the remaining mistakes on imagenet. In NeurIPS, 2022. 2
  73. 73.Andreas Veit, Serge Belongie, and Theofanis Karaletsos. Conditional similarity networks. In CVPR, 2017. 1, 2, 3
  74. 74.Nam Vo, Lu Jiang, Chen Sun, Kevin Murphy, Li-Jia Li, Li Fei-Fei, and James Hays. Composing text and image for image retrieval-an empirical odyssey. In CVPR, 2019. 2, 7
  75. 75.Jiang Wang, Yang Song, Thomas Leung, Chuck Rosenberg, Jingbin Wang, James Philbin, Bo Chen, and Ying Wu. Learning fine-grained image similarity with deep ranking. In CVPR, June 2014. 1
  76. 76.Hui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah, Steven Rennie, Kristen Grauman, and Rogerio Feris. The fashion iq dataset: Retrieving images by combining side information and relative natural language feedback. CVPR, 2021. 2, 7
  77. 77.Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang, Lei Li, Weiwei Sun, and Wei-Ying Ma. Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations. In CVPR, 2019. 5
  78. 78.Tete Xiao, Xiaolong Wang, Alexei A Efros, and Trevor Darrell. What should not be contrastive in contrastive learning. ICLR, 2021. 4
  79. 79.Yahui Xu, Yi Bin, Guoqing Wang, and Yang Yang. Hierarchical composition learning for composed query image retrieval. In ACM Multimedia Asia, 2021. 7
  80. 80.Yujie Zhong, Relja Arandjelović, and Andrew Zisserman. Faces in places: Compound query retrieval. In BMVC, 2016. 7
  81. 81.Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krähenbühl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. In ECCV, 2022. 6, 7, 3

Citation

MLA
Vaze, S., et al. “GeneCIS: A Benchmark for General Conditional Image Similarity”. arXiv, 2023, http://arxiv.org/abs/2306.07969v1.
APA
Vaze, S., Carion, N., & Misra, I. (2023). GeneCIS: A Benchmark for General Conditional Image Similarity. arXiv. http://arxiv.org/abs/2306.07969v1
Chicago
Vaze, S., N. Carion, and I. Misra. 2023. “GeneCIS: A Benchmark for General Conditional Image Similarity”. arXiv. http://arxiv.org/abs/2306.07969v1.
Harvard
Vaze, S., Carion, N. and Misra, I. (2023) “GeneCIS: A Benchmark for General Conditional Image Similarity”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.07969v1.
Vancouver
1. Vaze S, Carion N, Misra I (2023) GeneCIS: A Benchmark for General Conditional Image Similarity. arXiv

BibTeX

@article{vaze2023genecis,
  title = {GeneCIS: A Benchmark for General Conditional Image Similarity},
  author = {Vaze, Sagar and Carion, Nicolas and Misra, Ishan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.07969v1},
  eprint = {2306.07969}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE