keyword
image embeddings
Image embeddings are dense numerical vectors that represent the visual content and semantic features of an image in a continuous mathematical space. Produced by deep learning architectures such as convolutional neural networks or vision transformers, these representations compress high-dimensional raw pixel data while preserving essential attributes such as colors, shapes, textures, and high-level conceptual meaning. Because visually or semantically similar images are mapped to nearby coordinates within the embedding space, algorithms can compare them using vector distance metrics. This capability makes image embeddings foundational for numerous computer vision and multimodal artificial intelligence tasks, including reverse image search, content-based recommendation, image classification, and the alignment of visual data with text, audio, or other modalities for cross-modal retrieval and generation.
3 items

A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image Models
James Urquhart Allingham, Jie Ren, Michael W. Dusenberry, Xiuye Gu, Yin Cui, Dustin Tran, Jeremiah Zhe Liu, Balaji Lakshminarayanan
Why you should read this
Proposes a bias-corrected zero-shot prompt weighting algorithm that automatically scores and ensembles prompts for text-image models without needing labeled validation data or manual prompt engineering.
Added
2026-10-03

Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, Radu Soricut
Why you should read this
Introduces Conceptual Captions, a 3.3-million-example dataset harvested and hypernymed from web alt-text, and demonstrates that training vision-language models on this large-scale data substantially reduces object hallucinations and improves open-domain image description quality.
We present a new dataset of image caption annotations, Conceptual Captions, which contains an order of magnitude more images than the MS-COCO dataset (Lin et al., 2014) and represents a wider variety of both images and image caption styles. We achieve this by extracting and filtering image caption annotations from billions of webpages. We also present quantitative evaluations of a number of image captioning models and show that a model architecture based on Inception-ResNet-v2 (Szegedy et al., 2016) for image-feature extraction and Transformer (Vaswani et al., 2017) for sequence modeling achieves the best performance when trained on the Conceptual Captions dataset.
Added
2026-09-13

ImageBind One Embedding Space to Bind Them All
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, Ishan Misra
Why you should read this
Demonstrates that aligning various modalities (audio, depth, thermal) to images automatically aligns them to each other, creating a universal embedding space.
We present ImageBind, an approach to learn a joint embedding across six different modalities - images, text, audio, depth, thermal, and IMU data. We show that all combinations of paired data are not necessary to train such a joint embedding, and only image-paired data is sufficient to bind the modalities together. ImageBind can leverage recent large scale vision-language models, and extends their zero-shot capabilities to new modalities just by using their natural pairing with images. It enables novel emergent applications 'out-of-the-box' including cross-modal retrieval, composing modalities with arithmetic, cross-modal detection and generation. The emergent capabilities improve with the strength of the image encoder and we set a new state-of-the-art on emergent zero-shot recognition tasks across modalities, outperforming specialist supervised models. Finally, we show strong few-shot recognition results outperforming prior work, and that ImageBind serves as a new way to evaluate vision models for visual and non-visual tasks.
Added
2026-01-28
