What If We Recaption Billions of Web Images with LLaMA-3?

Xianhang LiHaoqin TuMude HuiZeyu WangBingchen ZhaoJunfei XiaoSucheng RenJieru MeiQing LiuHuangjie Zheng

article2025ICML86 citations

Presents an open-source pipeline using a LLaMA-3-powered LLaVA model to recaption 1.3 billion web images in DataComp-1B, substantially improving downstream performance for both CLIP retrieval and text-to-image generation.

Listen

Modern vision-language artificial intelligence systems, such as image retrieval engines and text-to-image generators, depend heavily on billions of image-text pairs scraped from the internet. However, this raw web-crawled data is inherently noisy, frequently containing brief, low-quality descriptions that misalign with actual image contents. While proprietary systems have improved performance by regenerating descriptive captions at scale, these high-performing datasets and pipelines have remained predominantly closed to the broader open-source community due to extreme monetary and computational costs.

The article demonstrates an open-source pipeline to regenerate rich, descriptive captions for approximately 1.3 billion images from the public DataComp-1B dataset. It systematically evaluates how training both discriminative models (which match text to images) and generative models (which create images from text) on this enhanced dataset impacts overall performance and text understanding.

To achieve this, the authors built an automated captioning model by integrating the open-source LLaMA-3 language model with the LLaVA vision-language architecture. After fine-tuning this captioner, they recaptioned the entire DataComp-1B dataset, expanding average text lengths from roughly 10 words to nearly 50 words while substantially diversifying vocabulary. They then trained and evaluated various configurations of dual-encoder retrieval models and diffusion-based image generators using varying blends of original and synthetic captions.

The findings confirm that enriched synthetic descriptions significantly enhance multimodal performance. For cross-modal retrieval models, mixing generated captions with original data produced an average 3.1% boost across standard benchmarks, with long-text retrieval improving by up to 36% and fine-grained attribute comprehension rising by over 6% to 9%. Text-to-image generative models trained on the recaptioned data exhibited marked improvements in image quality and prompt alignment, reducing image error scores by 8.4 points and raising alignment ratings across automated and human reviews. The evaluations also revealed that while purely synthetic captions degrade basic image classification, blending approximately 80% original captions with 20% generated captions preserves classification accuracy while capturing the full benefits of enhanced text retrieval.

These results show that descriptive synthetic data resolves significant data bottlenecks in vision-language pre-training, enabling open-source models to match or exceed the performance of models trained on vastly larger proprietary datasets. This substantially improves training efficiency and reduces computational overhead. However, practitioners must balance caption sources, as retaining short, original captions acts as a necessary regularizer against overfitting to synthetic text styles.

Organizations developing vision-language foundation models should adopt mixed-caption pre-training strategies rather than relying exclusively on raw web metadata or purely synthetic text. Future efforts should explore targeted classification-oriented recaptioning strategies, prompt conditioning on original metadata to capture specific entity names, and lightweight post-filtering to remove inherited algorithmic biases. Decision-makers should maintain moderate caution regarding lingering web-data safety risks, potential model hallucinations, and the licensing restrictions associated with foundational language model weights.

Cover for What If We Recaption Billions of Web Images with LLaMA-3?

Abstract

Large multimodal models trained at scale have demonstrated remarkable performance across diverse visual tasks.

Table of Contents

  • 1. Introduction
  • 2. Related works
  • 3. Recaptioning Pipeline
  • 3.1. Model details
  • 3.2. Recaptioning DataComp-1B
  • 4. Analyzing Recap-DataComp-1B
  • 4.1. Word & Length Distribution
  • 4.2. GPT-4V & CLIP Evaluations
  • 5. Training CLIPs with Recaptions
  • 5.1. Experiment settings
  • 5.2. Training with Mixed Captions
  • 5.3. Training with Larger Text Encoder
  • 5.4. More evaluations on text understanding
  • 5.5. Scaling-up Recap-CLIP
  • 6. Training Text-to-Image Models with Recaptions
  • 7. Conclusion
  • Acknowledgement
  • References
  • A. GPT-4V & Human Evaluation
  • B. Training cost
  • C. Limitation
  • D. License
  • E. DiT Qualitative Results
  • F. Ablations on Model and Prompt Selections

Knowls

  1. Knowl 1 — LLaMA-3-powered LLaVA captioner and its training

    model/method

    The captioner used to create Recap-DataComp-1B is a LLaVA-1.5-style model with a LLaMA-3-8B language decoder and a CLIP ViT-L/14 vision encoder. The vision encoder is frozen; two trainable MLP layers project its visual features into the language model's embedding space. Training uses autoregressive instruction tuning in two stages: first, train only the projection MLP on 558,000 image-text pairs filtered from LAION, Conceptual Captions, and SBU; second, train the MLP and language decoder on 665,000 LLaVA-1.5 instruction examples, including image-grounded conversation, descriptions, and visual reasoning. The authors also use HQ-Edit image-text pairs for further tuning to improve caption quality.

  2. Knowl 2 — Billion-scale construction of Recap-DataComp-1B

    model/method

    The authors apply the trained LLaVA-1.5-LLaMA3-8B captioner to approximately 1.3 billion web-crawled image-text pairs in DataComp-1B, a curated subset of a 12.8-billion-pair collection. For each image, the captioner receives the prompt “Please generate a detailed caption of this image. Please be as descriptive as possible.” It generates a caption autoregressively using greedy decoding, with a maximum output length of 128 tokens. The resulting image-caption dataset is named Recap-DataComp-1B.

  3. Knowl 3 — Recaptions are longer and score better on caption-quality evaluations

    empirical result

    Recap-DataComp-1B captions are substantially longer than the original DataComp-1B captions: their reported mean lengths are 49.43 and 10.22 tokens, respectively. In a randomly sampled analysis of about 0.35 billion pairs, the recaptions account for 82.86% of the tokens in the combined word collection, and the authors report more varied noun and adjective usage. On 180,000 image-caption pairs, standard CLIP scores are similar for original and recaptioned text (50.43 and 49.57, respectively), whereas LongCLIP-Large scores are 10.09 and 89.91. In GPT-4V ratings of 10,000 randomly selected examples on a 1–5 scale, recaptions average 4.14 versus 3.71 for original captions. A double-blind human rating of 200 randomly selected images likewise favors recaptions, with averages of 4.3 versus 3.1. These results support improved detail, fluency, and image-caption alignment, while not establishing that every generated detail is image-grounded.

  4. Knowl 4 — Mixing original and generated captions improves CLIP retrieval

    empirical result

    The authors train CLIP-B/16 with a randomly chosen caption per image: the original DataComp caption with probability pp, or the Recap-DataComp caption with probability 1−p1-p. Training follows a CLIPA-style two-stage schedule, using text length 128, 2.56 billion pretraining samples at image size 112, then 128 million fine-tuning samples at image size 224. The table reports zero-shot ImageNet-1K top-1 accuracy and retrieval Recall@1 (percent) on COCO and Flickr30K; I→T means image-to-text and T→I means text-to-image. The paper reports that retrieval generally improves over the original-caption baseline (p=1p=1), peaking around p=0.4p=0.4, while ImageNet classification declines when generated captions are introduced. The authors choose p=0.8p=0.8 for later experiments as a balance between retrieval gains and classification loss.

    pp ImageNet-1K COCO I→\toT COCO T→\toI Flickr30K I→\toT Flickr30K T→\toI
    0.0 36.0 53.0 34.1 74.1 53.5
    0.1 59.6 62.5 41.6 84.2 65.5
    0.2 63.8 61.7 42.4 86.8 67.0
    0.3 65.5 62.7 42.6 86.2 68.4
    0.4 66.7 63.4 43.2 87.6 68.2
    0.5 67.9 62.3 43.3 85.6 67.7
    0.6 68.8 63.6 43.1 86.2 68.2
    0.7 69.0 62.9 42.8 85.7 68.1
    0.8 69.8 62.8 42.7 86.7 67.1
    0.9 70.3 62.8 41.8 86.1 66.8
    1.0 70.5 59.5 38.9 84.1 64.4
  5. Knowl 5 — Recaption-trained CLIP scales to strong public-data retrieval results

    empirical result

    For larger-scale comparisons, the authors use a training recipe with p=0.8p=0.8 and a larger text encoder, scale training to 12.8 billion samples, and train L/14 and H/14 models. The reported metrics are ImageNet-1K top-1 accuracy and COCO/Flickr30K Recall@1 (percent). Recap-CLIP L/14 outperforms the listed public-data CLIP models on all five reported metrics, including the original DataComp-1B L/14 model. Recap-CLIP H/14 has stronger retrieval scores than the listed SigLIP SO(400M) model on COCO and Flickr30K image-to-text retrieval, but not on every metric. The authors report average retrieval gains over DataComp-1B L/14 of 5.6% on Flickr30K and 8.4% on COCO.

    Model Training data ImageNet-1K Flickr T→\toI Flickr I→\toT COCO T→\toI COCO I→\toT
    CLIP Large DataComp-1B 79.2 73.4 89.0 45.7 63.3
    SigLIP Large WebLI-5B 80.5 79.0 91.8 52.3 70.8
    Recap-CLIP L/14 Recap-DataComp-1B 79.3 79.5 94.1 53.7 72.0
    CLIP Huge DFN-5B 84.4 82.0 94.0 55.6 71.9
    SigLIP SO(400M) WebLI-5B 83.1 83.0 94.3 54.2 72.4
    Recap-CLIP H/14 Recap-DataComp-1B 81.0 81.3 94.8 54.5 73.1
  6. Knowl 6 — Larger CLIP text encoders retain recaptioning benefits

    empirical result

    With the original-caption probability fixed at p=0.8p=0.8, the authors compare CLIP models using larger text encoders while keeping the vision branch configuration fixed. Across S/16, B/16, and L/16 vision-model scales, recaption-based training improves retrieval Recall@1 over original-caption training by an average of 4.6%, 3.1%, and 3.3%, respectively. The reported ImageNet-1K top-1 accuracy is lower by 0.4–0.7 percentage points for recaption-trained counterparts. Enlarging the text encoder further improves retrieval performance, and the recaptioning benefit is observed across the tested model scales and text-encoder sizes.

  7. Knowl 7 — Recap-CLIP improves long-caption retrieval and attribute understanding

    empirical result

    The authors evaluate zero-shot text understanding on Urban1K, a 1,000-image urban retrieval benchmark with GPT-4V captions, and VG-Attribute, which measures attribute understanding. Scores are reported as percentages. Recap-CLIP trained with recaptions improves substantially over the same model trained with original captions: the B/16 Urban1K scores rise from 53.2 to 85.0 for I→T and from 50.9 to 87.3 for T→I, while its VG-Attribute score rises from 57.1 to 66.4. The L/16 model rises from 69.8 to 89.0, 64.6 to 91.8, and 60.1 to 66.8 on those respective measures. These results show better long-caption and attribute handling without a task-specific fine-tuning step.

  8. Knowl 8 — Recaption data improves text-to-image generation, especially with recaptioned prompts

    empirical result

    The authors train Diffusion Transformers (DiT) on DataComp-1B using mixtures with original-caption probability pp and recaption probability 1−p1-p. The DiT text condition comes from a CLIP text encoder; images are processed at 256-pixel square resolution and represented as latents from an autoencoder with downsampling ratio 8. Training uses batch size 2,048, AdamW, constant learning rate 10−410^{-4}, and no warm-up or weight decay. Evaluation generates 30,000 images with classifier-free guidance 10 and 250 DDPM steps. “Raw” evaluates against original COCO captions; “Our COCO-Recap” uses recaptioned COCO captions. FID is lower-is-better; the other reported scores are higher-is-better. GPT-4V scores use 3,000 generated images. When evaluated with recaptioned prompts, the all-recaption-trained model (p=0p=0) improves over the original-caption-trained model (p=1p=1) by 8.4 FID points, 3.1 CLIP-score points, 8.4 Recap-CLIP-score points, and 1.1 GPT-4V points. With raw prompts, the authors observe improved CLIP alignment but no consistent FID improvement. Their prose identifies p=0.1p=0.1 as the best overall mixture; the table gives the metric-specific values below.

    pp Raw FID Raw CLIP score Recap-prompt FID Recap-prompt CLIP score Recap-CLIP score GPT-4V score
    0.00 37.6 29.2 27.8 32.5 28.3 2.53
    0.05 38.5 29.1 27.9 32.5 28.0 2.51
    0.10 36.0 29.7 27.2 32.7 28.2 2.51
    0.15 35.8 29.9 28.2 33.0 28.1 2.45
    0.20 35.8 29.8 28.4 32.7 28.0 2.53
    0.50 35.3 29.3 30.2 31.9 26.7 2.13
    0.75 31.3 29.4 32.7 31.2 25.8 1.89
    1.00 32.5 28.9 36.2 29.3 19.9 1.40

    For a larger DiT-L/2 trained for one epoch at p=0p=0, the authors report FID 25.14 and CLIP score 34.82.

  9. Knowl 9 — LLaVA-1.5-LLaMA3-8B benchmark performance

    empirical result

    The captioner’s visual understanding and reasoning are evaluated on MMMU and MM-Vet. Its scores exceed those of LLaVA-1.5-7B on both benchmarks and are slightly above LLaVA-1.5-13B, despite using an 8B language decoder. GPT-4V scores are included as a reference.

    Model MMMU MM-Vet
    LLaVA-1.5-7B 33.6 33.9
    LLaVA-1.5-13B 36.4 36.3
    LLaVA-1.5-LLaMA3-8B (ours) 37.5 36.5
    GPT-4V 56.8 44.6
  10. Knowl 10 — Risks and limitations of the released recaption data

    limitation

    Recap-DataComp-1B inherits risks from its Common Crawl-derived source, including noisy or potentially unsafe content that may remain despite DataComp filtering, NSFW checks, face blurring, and a takedown policy. The captioner’s web-derived training data may also transmit gender, racial, or cultural biases, and its generated descriptions can include hallucinated details. The authors note that the model may fail to identify specific named entities, and that their CLIP experiments show weaker zero-shot classification when trained with recaptions, particularly when original captions are removed. They also state that computational and inference costs constrained their choice of captioning model.

Coverage note — Detailed prompt/model ablations, inference and training cost estimates, licensing terms, and illustrative caption examples are omitted because they are supporting or operational material rather than additional central findings.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023a.
  2. 2.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. GPT-4V(ision) system card. OpenAI Research Blog, 2023b.
  3. 3.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. In NeurIPS, 2022.
  4. 4.Awadalla, A., Xue, L., Lo, O., Shu, M., Lee, H., Guha, E. K., Jordan, M., Shen, S., Awadalla, M., Savarese, S., et al. Mint-1t: Scaling open-source multimodal data by 10x: A multimodal dataset with one trillion tokens. arXiv preprint arXiv:2406.11271, 2024.
  5. 5.Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv preprint arXiv:2308.12966, 2023.
  6. 6.Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2023.
  7. 7.Changpinyo, S., Sharma, P., Ding, N., and Soricut, R. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR, 2021.
  8. 8.Chen, J., Ge, C., Xie, E., Wu, Y., Yao, L., Ren, X., Wang, Z., Luo, P., Lu, H., and Li, Z. Pixart-σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation. arXiv preprint arXiv:2403.04692, 2024a.
  9. 9.Chen, J., YU, J., GE, C., Yao, L., Xie, E., Wang, Z., Kwok, J., Luo, P., Lu, H., and Li, Z. Pixart-α\alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis. In ICLR, 2024b.
  10. 10.Chen, L., Li, J., Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., and Lin, D. Sharegpt4v: Improving large multi-modal models with better captions. arXiv preprint arXiv:2311.12793, 2023a.
  11. 11.Chen, X. and Wang, X. Pali: Scaling language-image learning in 100+ languages. In NeurIPS, 2022.
  12. 12.Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., Li, B., Luo, P., Lu, T., Qiao, Y., and Dai, J. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. arXiv preprint arXiv:2312.14238, 2023b.
  13. 13.Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhang, H., Zhu, B., Jordan, M., Gonzalez, J. E., and Stoica, I. Chatbot arena: An open platform for evaluating llms by human preference. arXiv preprint arXiv:2403.04132, 2024.
  14. 14.Chu, X., Qiao, L., Zhang, X., Xu, S., Wei, F., Yang, Y., Sun, X., Hu, Y., Lin, X., Zhang, B., et al. Mobilevlm v2: Faster and stronger baseline for vision language model. arXiv preprint arXiv:2402.03766, 2024.
  15. 15.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  16. 16.Desai, K., Kaul, G., Aysola, Z., and Johnson, J. Redcaps: Web-curated image-text data created by the people, for the people. arXiv preprint arXiv:2111.11431, 2021.
  17. 17.Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al. Cogview: Mastering text-to-image generation via transformers. In NeurIPS, 2021.
  18. 18.Fan, L., Krishnan, D., Isola, P., Katabi, D., and Tian, Y. Improving clip training with language rewrites. In NeurIPS, 2024.
  19. 19.Fang, A., Jose, A. M., Jain, A., Schmidt, L., Toshev, A., and Shankar, V. Data filtering networks. arXiv preprint arXiv:2309.17425, 2023.
  20. 20.Fei, Z., Fan, M., Yu, C., Li, D., Zhang, Y., and Huang, J. Dimba: Transformer-mamba diffusion models. arXiv preprint arXiv:2406.01159, 2024.
  21. 21.Gadre, S. Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al. Datacomp: In search of the next generation of multimodal datasets. arXiv preprint arXiv:2304.14108, 2023.
  22. 22.Gemma Team. Gemma 2: Improving open language models at a practical size. ArXiv, abs/2408.00118, 2024.
  23. 23.Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NeurIPS, 2017.
  24. 24.Hinck, M., Olson, M. L., and Lal, V. Intel/llava-llama-3-8b: Llava-v1.5 fine-tuned meta-llama-3-8b multimodal model. URL https://huggingface.co/Intel/llava-llama-3-8b. Intel Research Use Licence. Accessed 2025-05-29.
  25. 25.Huang, H., Peng, S., Zhang, D., and Geiger, A. Renovating names in open-vocabulary segmentation benchmarks. arXiv preprint arXiv:2403.09593, 2024.
  26. 26.Hui, M., Yang, S., Zhao, B., Shi, Y., Wang, H., Wang, P., Zhou, Y., and Xie, C. Hq-edit: A high-quality dataset for instruction-based image editing. arXiv preprint arXiv:2404.09990, 2024.
  27. 27.Ilharco, G., Wortsman, M., Wightman, R., Gordon, C., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., and Schmidt, L. Openclip. github, July 2021. doi: 10.5281/zenodo.5143773. URL https://doi.org/10.5281/zenodo.5143773.
  28. 28.Kang, M., Zhu, J.-Y., Zhang, R., Park, J., Shechtman, E., Paris, S., and Park, T. Scaling up GANs for text-to-image synthesis. In CVPR, 2023.
  29. 29.Karpathy, A. and Fei-Fei, L. Deep visual-semantic alignments for generating image descriptions. In CVPR, 2015.
  30. 30.Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. In IJCV, 2017.
  31. 31.Lee, T., Mai, Y., Wong, C. H., Roberts, J. S., Yasunaga, M., Kaiyom, F., Bommasani, R., and Liang, P. The first steps to holistic evaluation of vision-language models, May 2024.
  32. 32.Li, J., Li, D., Xiong, C., and Hoi, S. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, 2022.
  33. 33.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2023a.
  34. 34.Li, X., Wang, Z., and Xie, C. An inverse scaling law for clip training. In NeurIPS, 2023b.
  35. 35.Li, X., Wang, Z., and Xie, C. Clipa-v2: Scaling clip training with 81.1arXiv preprint arXiv:2306.15658, 2023c.
  36. 36.Lin, B., Tang, Z., Ye, Y., Cui, J., Zhu, B., Jin, P., Huang, J., Zhang, J., Ning, M., and Yuan, L. Moe-llava: Mixture of experts for large vision-language models. arXiv preprint arXiv:2401.15947, 2024.
  37. 37.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pp. 740–755. Springer, 2014.
  38. 38.Liu, H., Li, C., Li, Y., and Lee, Y. J. Improved baselines with visual instruction tuning. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following, 2023a.
  39. 39.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023b.
  40. 40.Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J. Llava-next: Improved reasoning, ocr, and world knowledge. https://llava-vl.github.io/blog/2024-01-30-llava-next/, January 2024.
  41. 41.Liu, X., Zhang, X., Ma, J., Peng, J., and Liu, Q. Instaflow: One step is enough for high-quality diffusion-based text-to-image generation. arXiv preprint arXiv:2309.06380, 2023c.
  42. 42.Liu, Y., Wang, K., Shao, W., Luo, P., Qiao, Y., Shou, M. Z., Zhang, K., and You, Y. Mllms-augmented visual-language representation learning. arXiv preprint arXiv:2311.18765, 2023d.
  43. 43.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  44. 44.Lu, G., Guo, Y., Han, J., Niu, M., Zeng, Y., Xu, S., Huang, Z., Zhong, Z., Zhang, W., and Xu, H. Pangudraw: Advancing resource-efficient text-to-image synthesis with time-decoupled training and reusable coop-diffusion. arXiv preprint arXiv:2312.16486, 2023.
  45. 45.Meta LLaMA Team. Introducing Meta Llama 3: The most capable openly available LLM to date, 2024.
  46. 46.Nguyen, T., Gadre, S. Y., Ilharco, G., Oh, S., and Schmidt, L. Improving multimodal datasets with image captioning. In NeurIPS, 2024.
  47. 47.Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  48. 48.OpenAI. Dall·e 3 system card. OpenAI Research Blog, 2023.
  49. 49.OpenAI. Video generation models as world simulators. OpenAI Research Blog, 2024.
  50. 50.Ordonez, V., Kulkarni, G., and Berg, T. Im2text: Describing images using 1 million captioned photographs. In NeurIPS, 2011.
  51. 51.Padlewski, P., Bain, M., Henderson, M., Zhu, Z., Relan, N., Pham, H., Ong, D., Aleksiev, K., Ormazabal, A., Phua, S., et al. Vibe-eval: A hard evaluation suite for measuring progress of multimodal language models. arXiv preprint arXiv:2405.02287, 2024.
  52. 52.Peebles, W. and Xie, S. Scalable diffusion models with transformers. In ICCV, 2023.
  53. 53.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 8748–8763. PMLR, 18–24 Jul 2021a. URL https://proceedings.mlr.press/v139/radford21a.html.
  54. 54.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In ICML, 2021b.
  55. 55.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. In International Conference on Machine Learning, pp. 8821–8831. PMLR, 2021.
  56. 56.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents, 2022. URL https://arxiv.org/abs/2204.06125.
  57. 57.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models, 2021.
  58. 58.Rotstein, N., Bensaid, D., Brody, S., Ganz, R., and Kimmel, R. Fusecap: Leveraging large language models to fuse visual data into enriched image captions. arXiv preprint arXiv:2305.17718, 2023.
  59. 59.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015. doi: 10.1007/s11263-015-0816-y.
  60. 60.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al. Photorealistic text-to-image diffusion models with deep language understanding. In NeurIPS, 2022.
  61. 61.Sauer, A., Karras, T., Laine, S., Geiger, A., and Aila, T. StyleGAN-T: Unlocking the power of GANs for fast large-scale text-to-image synthesis. arXiv preprint arXiv:2301.09515, 2023a.
  62. 62.Sauer, A., Lorenz, D., Blattmann, A., and Rombach, R. Adversarial diffusion distillation. ArXiv, abs/2311.17042, 2023b.
  63. 63.Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021.
  64. 64.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kundurthy, S., Crowson, K., Schmidt, L., Kaczmarczyk, R., and Jitsev, J. Laion-5b: An open large-scale dataset for training next generation image-text models, 2022a.
  65. 65.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35:25278–25294, 2022b.
  66. 66.Sharma, P., Ding, N., Goodman, S., and Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL, 2018.
  67. 67.Srinivasan, K., Raman, K., Chen, J., Bendersky, M., and Najork, M. Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning. In SIGIR, 2021.
  68. 68.Stevens, I. Llama 3’s performance benchmark values explained, 2024. Accessed: 2024-06-05.
  69. 69.Sun, Z., Shen, S., Cao, S., Liu, H., Li, C., Shen, Y., Gan, C., Gui, L.-Y., Wang, Y.-X., Yang, Y., et al. Aligning large multimodal models with factually augmented rlhf. arXiv preprint arXiv:2309.14525, 2023.
  70. 70.Wang, Z., Yu, J., Yu, A. W., Dai, Z., Tsvetkov, Y., and Cao, Y. Simvlm: Simple visual language model pretraining with weak supervision. In ICLR, 2022.
  71. 71.Xu, H., Xie, S., Tan, X., Huang, P.-Y., Howes, R., Sharma, V., Li, S.-W., Ghosh, G., Zettlemoyer, L., and Feichtenhofer, C. Demystifying clip data. In ICLR, 2023.
  72. 72.Xu, R., Yao, Y., Guo, Z., Cui, J., Ni, Z., Ge, C., Chua, T.-S., Liu, Z., and Huang, G. LLaVA-UHD: an lmm perceiving any aspect ratio and high-resolution images. arXiv preprint arXiv:2403.11703, 2024.
  73. 73.Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. In TACL, 2014.
  74. 74.Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., Hutchinson, B., Han, W., Parekh, Z., Li, X., Zhang, H., Baldridge, J., and Wu, Y. Scaling autoregressive models for content-rich text-to-image generation, 2022a.
  75. 75.Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al. Scaling autoregressive models for content-rich text-to-image generation. In TMLR, 2022b.
  76. 76.Yu, Q., Sun, Q., Zhang, X., Cui, Y., Zhang, F., Wang, X., and Liu, J. Capsfusion: Rethinking image-text data at scale. arXiv preprint arXiv:2310.20550, 2023a.
  77. 77.Yu, T., Yao, Y., Zhang, H., He, T., Han, Y., Cui, G., Hu, J., Liu, Z., Zheng, H.-T., Sun, M., et al. Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback. arXiv preprint arXiv:2312.00849, 2023b.
  78. 78.Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L. Mm-vet: Evaluating large multimodal models for integrated capabilities. In ICML, 2024.
  79. 79.Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun, H., Su, Y., and Chen, W. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In CVPR, 2024.
  80. 80.Yuksekgonul, M., Bianchi, F., Kalluri, P., Jurafsky, D., and Zou, J. When and why vision-language models behave like bags-of-words, and what to do about it? In ICLR, 2022.
  81. 81.Zhai, X., Wang, X., Mustafa, B., Steiner, A., Keysers, D., Kolesnikov, A., and Beyer, L. Lit: Zero-shot transfer with locked-image text tuning. In CVPR, 2022.
  82. 82.Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L. Sigmoid loss for language image pre-training. In ICCV, 2023.
  83. 83.Zhang, B., Zhang, P., Dong, X., Zang, Y., and Wang, J. Long-clip: Unlocking the long-text capability of clip. arXiv preprint arXiv:2403.15378, 2024a.
  84. 84.Zhang, L., Shu, F., Ren, S., Zhao, B., Jiang, H., and Xie, C. Compress & align: Curating image-text data with human knowledge. arXiv preprint arXiv:2312.06726, 2023.
  85. 85.Zhang, Y., Unell, A., Wang, X., Ghosh, D., Su, Y., Schmidt, L., and Yeung-Levy, S. Why are visually-grounded language models bad at image classification? arXiv preprint arXiv:2405.18415, 2024b.
  86. 86.Zheng, K., Zhang, Y., Wu, W., Lu, F., Ma, S., Jin, X., Chen, W., and Shen, Y. Dreamlip: Language-image pre-training with long captions. In ECCV, 2024.
  87. 87.Zhou, M., Wang, Z., Zheng, H., and Huang, H. Long and short guidance in score identity distillation for one-step text-to-image generation. arXiv preprint arXiv:2406.01561, 2024.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/