Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
Piyush SharmaNan DingSebastian GoodmanRadu Soricut
Introduces Conceptual Captions, a 3.3-million-example dataset harvested and hypernymed from web alt-text, and demonstrates that training vision-language models on this large-scale data substantially reduces object hallucinations and improves open-domain image description quality.
A new dataset called Conceptual Captions addresses the limited size and variety of existing resources for training automatic image captioning systems. Current collections such as MS-COCO contain roughly 120,000 images with human-written captions and favor a narrow range of everyday scenes, which restricts model generalization and encourages hallucinations of objects that are statistically common in the training data rather than present in a given image. The work extracts candidate image-description pairs at web scale from alt-text attributes, applies successive filters for image quality, text well-formedness, and visual-textual overlap, then replaces specific names and details with appropriate hypernyms to produce clean, learnable captions.
The authors constructed a pipeline that processed billions of webpages and retained approximately 3.3 million image-caption pairs after rigorous filtering and transformation steps. They trained both RNN-based and Transformer-based captioning models on this collection, using an Inception-ResNet-v2 network to extract image features, and compared performance against identical models trained on COCO data. Evaluation covered in-domain and out-of-domain test sets, including the Flickr 1K set, with both automatic metrics and human ratings from professional annotators.
Human evaluations on the out-of-domain Flickr test set showed that Transformer models trained on Conceptual Captions received majority “good” ratings in 50.6 percent of cases, compared with 36.2 percent for the same architecture trained on COCO. These models also produced fewer hallucinations, handled cartoons and product images without collapse, and generated more precise terms such as “graduates” or “cloister of the cathedral.” Automatic metrics confirmed strong in-domain results for Conceptual Captions models but diverged from human judgments on out-of-domain data, underscoring known weaknesses in current evaluation measures.
The findings indicate that larger, web-harvested datasets with controlled generalization can materially improve caption quality and robustness without requiring manual annotation. Practical systems for image retrieval or accessibility tools would therefore benefit from training on Conceptual Captions or similar resources. The authors recommend further exploration of Transformer architectures, which deliver higher accuracy with lower computational cost than RNN alternatives.
The main limitations are the low overall yield of the filtering pipeline (0.2 percent of candidates) and the reliance on automatic metrics that under-penalize hallucinations. Readers should treat reported automatic scores with caution when comparing across domains.
- Paper: Microsoft COCO Captions: Data Collection and Evaluation Server, Xinlei Chen et al. (2015). This paper establishes the MS-COCO Captions benchmark and standard evaluation protocols that Conceptual Captions explicitly contrasts against and uses as a baseline.
- Paper: Im2Text: Describing Images Using 1 Million Captioned Photographs, Vicente Ordonez et al. (2011). This work pioneered harvesting and filtering large-scale, web-sourced photograph-caption pairs from user data, providing direct intellectual foundation for web-scale alt-text dataset creation.
- Paper: Show and tell: A neural image caption generator, Oriol Vinyals et al. (2015). This foundational paper introduces the standard CNN-RNN encoder-decoder architecture for image captioning that serves as the baseline modeling approach evaluated on Conceptual Captions.
- Paper: Show, Attend and Tell: Neural Image Caption Generation with Visual Attention, Kelvin Xu et al. (2015). This paper establishes the attention mechanism in neural image captioning that motivates the sequence modeling architectures analyzed in Conceptual Captions.
- Paper: CIDEr: Consensus-based image description evaluation, Ramakrishna Vedantam et al. (2014). This work introduces the CIDEr evaluation metric used to benchmark and evaluate image captioning models throughout the Conceptual Captions experiments.
- Paper: SPICE: Semantic Propositional Image Caption Evaluation, Peter Anderson et al. (2016). This study introduces semantic propositional evaluation (SPICE) for image descriptions, informing the evaluation critique and metrics discussed in Conceptual Captions.
- Paper: Deep visual-semantic alignments for generating image descriptions, Andrej Karpathy et al. (2015). This work demonstrates deep multimodal visual-semantic alignment for caption generation and retrieval, which forms the basis for cross-modal filtering pipelines.
- Paper: Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts, Soravit Changpinyo et al. (2021). This direct follow-up scales Conceptual Captions to Conceptual 12M by relaxing the strict filtering rules to improve coverage of long-tail visual concepts.
- Paper: Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision, Chao Jia et al. (2021). This work scales web-scraped alt-text training to over one billion pairs with minimal filtering, testing the limits of scale versus curation compared to Conceptual Captions.
- Paper: LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs, Christoph Schuhmann et al. (2021). This paper advances web-scale dataset curation by using multimodal embeddings to filter hundreds of millions of image-text pairs into an open benchmark.
- Paper: CoCa: Contrastive Captioners are Image-Text Foundation Models, Jiahui Yu et al. (2022). This work builds foundation models that combine contrastive pretraining with autoregressive captioning on large-scale web alt-text data.
- Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). This paper utilizes large web-scraped image-text datasets, including Conceptual Captions, to achieve zero-shot autoregressive text-to-image generation.
- Paper: CLIPScore: A Reference-free Evaluation Metric for Image Captioning, Jack Hessel et al. (2021). This study develops a reference-free evaluation metric using vision-language embeddings, directly addressing the evaluation limitations and metric disconnects identified in Conceptual Captions.
- Paper: DataComp: In search of the next generation of multimodal datasets, Samir Yitzhak Gadre et al. (2023). This benchmark systematicizes the evaluation of data filtering and curation strategies on web-harvested multimodal pools pioneered by Conceptual Captions.
- Paper: Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling, Dat Huynh et al. (2022). This paper applies Conceptual Captions paired data to perform open-vocabulary instance segmentation without manual mask annotations.
