GeneCIS: A Benchmark for General Conditional Image Similarity
Sagar VazeNicolas CarionIshan Misra
Introduces the GeneCIS benchmark to evaluate zero-shot conditional image similarity and presents a scalable training strategy that mines image-caption datasets to significantly improve retrieval models when adapting to open-ended similarity criteria.
Standard computer vision models typically rely on a single, fixed notion of image similarity, such as matching overall object categories. However, practical human workflows require dynamic similarity comparisons based on context, such as identifying items with the same color, isolating a specific object in a complex scene, or modifying a single attribute. The article addresses this limitation by formalizing the task of general conditional image similarity, where models must retrieve relevant images based on an explicit user-specified text prompt without task-specific retraining.
The main objective of the article is to establish a rigorous evaluation benchmark for conditional image similarity and introduce a scalable, automated training approach that enables models to adapt to diverse visual conditions in an open-set, zero-shot setting.
To accomplish this, the authors introduced the GeneCIS benchmark, comprising four distinct evaluation tasks constructed from existing visual datasets: focusing on an attribute, changing an attribute, focusing on an object, and changing an object. Each task evaluates a model's ability to select the correct target image from small galleries containing challenging distractor images. To train models without labor-intensive manual labeling, the authors developed an automated pipeline that mines 1.6 million training triplets from 3 million web image-caption pairs. This method parses captions into subject-predicate-object relationships, filters them for visual concreteness using an external linguistic database, and pairs images sharing common subjects under modified conditions using a contrastive learning architecture.
The investigation produced several key findings. First, established vision-language models struggle significantly on conditional retrieval; simple image-and-text averaging achieved only a 12.6% average top-1 recall across the GeneCIS tasks. Second, performance on conditional similarity is only weakly correlated with standard benchmark accuracy: a 10% gain in zero-shot ImageNet accuracy yielded approximately a 1% gain on GeneCIS. Third, the proposed automated caption-mining approach achieved a 16.8% average top-1 recall, outperforming all zero-shot baselines as well as models trained on manually curated datasets. Finally, when tested on external composed image retrieval benchmarks, the zero-shot model surpassed supervised state-of-the-art models on the MIT-States benchmark (15.8% vs. 15.6% top-1 recall) and outperformed zero-shot baselines on the CIRR benchmark (27.3% vs. 21.8% top-1 recall).
These findings demonstrate that general vision systems cannot rely solely on standard scaling to solve complex, instruction-based retrieval tasks. Instead, explicit conditioning mechanisms are required. By leveraging existing, abundant image-caption data, organizations can significantly enhance multi-modal search and interactive computer vision applications without the high costs and timelines associated with bespoke manual data annotation.
Organizations developing fine-grained visual search or interactive media systems should adopt automated relationship parsing on large caption datasets rather than investing in manual condition labeling. Practitioners should also evaluate their vision backbones directly on instruction-based benchmarks rather than relying on general classification metrics to predict conditional retrieval performance. Further work should explore scaling this automated mining technique to web-scale datasets containing billions of image-text pairs.
Confidence in these findings is supported by consistent gains across multiple distinct benchmarks and ablation studies. However, the initial GeneCIS release (version 0) contains minor underlying label noise inherited from source datasets, and the current training triplet distribution exhibits a natural bias toward object-modification tasks rather than attribute adjustments. Users should account for these boundary conditions when deploying the approach across specialized domains.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. It introduces the CLIP contrastive vision-language pre-training framework that forms the fundamental baseline and representation backbone evaluated and modified in GeneCIS.
- Paper: Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts, Soravit Changpinyo et al. (2021). It details scaling web-scraped image-caption datasets for pre-training vision-language models, providing the empirical foundation for automated caption mining.
- Paper: Conditional Prompt Learning for Vision-Language Models, Kaiyang Zhou et al. (2022). It demonstrates instance-conditional prompt learning for vision-language models, establishing the value of input-dependent visual conditioning.
- Paper: Learning Fine-Grained Image Similarity with Deep Ranking, Jiang Wang et al. (2014). It formalizes deep triplet ranking for fine-grained visual similarity, which directly underlies the contrastive triplet formulation used to train conditional similarity.
- Paper: Image Difference Captioning with Pre-training and Contrastive Learning, Linli Yao et al. (2022). It presents methods for aligning subtle visual distinctions with language descriptions, establishing essential concepts for evaluating image differences.
- Paper: CoCa: Contrastive Captioners are Image-Text Foundation Models, Jiahui Yu et al. (2022). It unifies contrastive cross-modal representation learning with captioning decoders, providing foundational multi-modal pre-training methodologies.
- Paper: Stacked Cross Attention for Image-Text Matching, Kuang-Huei Lee et al. (2018). It introduces stacked cross-attention to match fine-grained image regions with text descriptions, motivating latent attribute-level visual alignment.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). It builds on conditional visual evaluation by introducing multimodal LLM-driven metrics to score semantic consistency in conditional image synthesis.
- Paper: PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers, Weizhe Lin et al. (2024). It scales multi-modal retrieval by employing late-interaction architectures over complex, query-conditioned image-text tasks.
- Paper: Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID, Wentao Tan et al. (2024). It leverages multimodal LLMs to automatically generate conditioned visual retrieval datasets while handling noisy captions.
- Paper: SoftCLIP: Softer Cross-Modal Alignment Makes CLIP Stronger, Yuting Gao et al. (2024). It refines contrastive vision-language pre-training by softening rigid one-to-one alignments to better capture nuanced, overlapping visual relationships.
- Paper: T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation, Kaiyue Sun 0001 et al. (2025). It extends compositional evaluation paradigms from static conditional image settings to dynamic, spatio-temporal video generation.
