What If We Recaption Billions of Web Images with LLaMA-3?
Xianhang LiHaoqin TuMude HuiZeyu WangBingchen ZhaoJunfei XiaoSucheng RenJieru MeiQing LiuHuangjie Zheng
Presents an open-source pipeline using a LLaMA-3-powered LLaVA model to recaption 1.3 billion web images in DataComp-1B, substantially improving downstream performance for both CLIP retrieval and text-to-image generation.
Modern vision-language artificial intelligence systems, such as image retrieval engines and text-to-image generators, depend heavily on billions of image-text pairs scraped from the internet. However, this raw web-crawled data is inherently noisy, frequently containing brief, low-quality descriptions that misalign with actual image contents. While proprietary systems have improved performance by regenerating descriptive captions at scale, these high-performing datasets and pipelines have remained predominantly closed to the broader open-source community due to extreme monetary and computational costs.
The article demonstrates an open-source pipeline to regenerate rich, descriptive captions for approximately 1.3 billion images from the public DataComp-1B dataset. It systematically evaluates how training both discriminative models (which match text to images) and generative models (which create images from text) on this enhanced dataset impacts overall performance and text understanding.
To achieve this, the authors built an automated captioning model by integrating the open-source LLaMA-3 language model with the LLaVA vision-language architecture. After fine-tuning this captioner, they recaptioned the entire DataComp-1B dataset, expanding average text lengths from roughly 10 words to nearly 50 words while substantially diversifying vocabulary. They then trained and evaluated various configurations of dual-encoder retrieval models and diffusion-based image generators using varying blends of original and synthetic captions.
The findings confirm that enriched synthetic descriptions significantly enhance multimodal performance. For cross-modal retrieval models, mixing generated captions with original data produced an average 3.1% boost across standard benchmarks, with long-text retrieval improving by up to 36% and fine-grained attribute comprehension rising by over 6% to 9%. Text-to-image generative models trained on the recaptioned data exhibited marked improvements in image quality and prompt alignment, reducing image error scores by 8.4 points and raising alignment ratings across automated and human reviews. The evaluations also revealed that while purely synthetic captions degrade basic image classification, blending approximately 80% original captions with 20% generated captions preserves classification accuracy while capturing the full benefits of enhanced text retrieval.
These results show that descriptive synthetic data resolves significant data bottlenecks in vision-language pre-training, enabling open-source models to match or exceed the performance of models trained on vastly larger proprietary datasets. This substantially improves training efficiency and reduces computational overhead. However, practitioners must balance caption sources, as retaining short, original captions acts as a necessary regularizer against overfitting to synthetic text styles.
Organizations developing vision-language foundation models should adopt mixed-caption pre-training strategies rather than relying exclusively on raw web metadata or purely synthetic text. Future efforts should explore targeted classification-oriented recaptioning strategies, prompt conditioning on original metadata to capture specific entity names, and lightweight post-filtering to remove inherited algorithmic biases. Decision-makers should maintain moderate caution regarding lingering web-data safety risks, potential model hallucinations, and the licensing restrictions associated with foundational language model weights.
- Paper: DataComp: In search of the next generation of multimodal datasets, Samir Yitzhak Gadre et al. (2023). Read DataComp first to understand the benchmark, CommonPool, and dataset-curation framework that supplies the source paper’s DataComp-1B data.
- Paper: Improving CLIP Training with Language Rewrites, Lijie Fan et al. (2023). LaCLIP establishes that language-model rewrites can improve CLIP training, providing a direct precedent for the source’s use of synthetic captions in image-text pretraining.
- Paper: BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, Junnan Li et al. (2022). BLIP’s captioning-and-filtering approach introduces the synthetic-caption strategy that helps frame the source’s recaptioning pipeline and mixed-caption experiments.
- Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). DALL·E shows how autoregressive image generation depends on large-scale web image-text training, clarifying one generative-model setting evaluated by the source.
- Paper: Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding, Chitwan Saharia et al. (2022). Imagen connects prompt understanding and diffusion-based synthesis, preparing readers for the source’s evaluation of recaptioned data on image generation.
- Paper: Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision, Chao Jia et al. (2021). ALIGN demonstrates contrastive vision-language learning at web scale with noisy text supervision, the training context the source seeks to improve through richer captions.
- Paper: LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs, Christoph Schuhmann et al. (2021). LAION-400M explains how open web image-text pairs are filtered and released for multimodal training, a key precedent for the source’s open-data motivation.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). InternVL3 carries the open multimodal-training agenda forward, testing newer integrated training and inference recipes after the source’s study of richer pretraining captions.
- Paper: Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward, Ruohong Zhang et al. (2025). This video-model work extends caption-based supervision into preference optimization, using detailed generated descriptions as scalable feedback for multimodal training.
