keyword
Conceptual Captions datasets
Conceptual Captions datasets are large-scale multimodal benchmarks consisting of millions of paired images and descriptive text captions designed for training and evaluating vision-and-language machine learning models. Unlike traditional caption datasets built through manual human labeling, these datasets are automatically harvested from web pages by extracting online images along with their associated alternative text descriptions. The collected pairs undergo automated filtering and text transformation pipelines that eliminate noisy or non-descriptive text, verify visual-semantic relevance, and replace specific proper names and fine-grained entities with broader conceptual categories. Widely used versions, such as Conceptual Captions 3M and Conceptual 12M, provide high-scale linguistic and visual diversity to support pre-training for tasks including image captioning, cross-modal retrieval, and open-vocabulary visual recognition.
1 item

