OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Hugo LaurençonLucile SaulnierLéo TronchonStas BekmanAmanpreet SinghAnton LozhkovThomas WangSiddharth KaramchetiAlexander M. RushDouwe Kiela

article2023NeurIPS342 citations

Introduces the first fully open, web-scale dataset of interleaved image-text documents alongside competitive 9B and 80B multimodal models, enabling reproducible research and training for next-generation vision-language systems.

Listen

Modern vision and language artificial intelligence models achieve superior capabilities when trained on naturally structured web documents that interleave text and images, rather than on simple image-caption pairs. However, the large-scale multimodal web datasets previously used to train leading proprietary systems have remained private, limiting research replication, transparency, and broader development in the artificial intelligence community.

The article demonstrates the viability of creating an open, high-quality, web-scale multimodal document dataset, named OBELICS, and evaluates its effectiveness by training open vision and language models named IDEFICS at both 9-billion and 80-billion parameter scales.

The authors constructed the dataset from 25 raw web dumps spanning February 2020 to February 2023, originally containing 41.2 billion documents. The pipeline filtered out non-English content, stripped boilerplate web elements using webpage document structure rules, removed duplicate text and images, filtered out adult content, and respected creator opt-out preferences. The final dataset yielded 141 million multimodal documents, 353 million images, and 115 billion text tokens. To validate the resource, the authors trained multimodal models using this data combined with open image-text pairs and benchmarked their performance across eight standard evaluation tasks.

The analysis reveals several critical findings. First, training on interleaved web documents achieves equivalent multimodal performance using an order of magnitude fewer images compared to training solely on isolated image-text pairs. Second, the 80-billion parameter IDEFICS model matches or exceeds the performance of comparable leading proprietary models across multiple visual reasoning benchmarks. Third, at the 9-billion parameter scale, the model outperforms competing open-source alternatives trained on older, less filtered web corpuses. Finally, the textual content in OBELICS exhibits significantly higher quality and lower perplexity scores than existing open multimodal and general web baselines.

These findings indicate that organizations can train highly competitive multimodal systems using curated, openly accessible web documents without relying on proprietary training data. Leveraging structured, interleaved documents also lowers the data volume requirements for visual pre-training, reducing computational overhead and infrastructure costs while preserving document context.

Teams developing multimodal models should adopt curated interleaved document corpuses, ideally combining them with image-text pairs to balance visual question answering and fine-grained classification capabilities. Before deploying models broadly, practitioners should implement domain-specific guardrails, as residual web crawling risks like occasional uncaptured advertisements or public imagery remain present. Researchers should continue developing finer semantic filtering and multi-language extensions to expand dataset utility.

Cover for OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Abstract

Large multimodal models trained on natural documents, which interleave images and text, outperform models trained on image-text pairs on various multimodal benchmarks. However, the datasets used to train these models have not been released, and the collection process has not been fully specified. We introduce the OBELICS dataset, an open web-scale filtered dataset of interleaved image-text documents comprising 141 million web pages extracted from Common Crawl, 353 million associated images, and 115 billion text tokens. We describe the dataset creation process, present comprehensive filtering rules, and provide an analysis of the dataset’s content. To show the viability of OBELICS, we train vision and language models of 9 and 80 billion parameters named IDEFICS, and obtain competitive performance on different multimodal benchmarks. We release our dataset, models and code.¹.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Creation of the Multimodal Web Document Dataset
  • 3.1 Collecting a Large Number of HTML Files
  • 3.2 Simplifying HTML Files
  • 3.3 Extracting Multimodal Web Documents
  • 3.4 Filtering Multimodal Web Documents
  • 3.5 Responsible Filtering and Deduplication
  • 4 Analysis of OBELICS
  • 4.1 General Statistics
  • 4.2 Topic Modeling
  • 4.3 Qualitative Assessment of Dataset Samples
  • 5 Validating the Viability of OBELICS
  • 6 Conclusion
  • Acknowledgments and Disclosure of Funding
  • References
  • Checklist

Knowls

  1. Knowl 1 — OBELICS Dataset Specification

    definition

    OBELICS is an open, web-scale, filtered dataset of naturally occurring interleaved image-text web documents extracted from 25 Common Crawl dumps spanning February 2020 to January/February 2023.

    The final curated dataset consists of:

    • Documents: 141 million multimodal English web pages.
    • Images: 353 million associated image instances, comprising 298 million unique images (an 84.3% uniqueness rate).
    • Tokens: 115 billion text tokens.
    • Document Structure: The dataset preserves the original DOM tree sequence, retaining line breaks, line feeds, and formatting to reflect the natural spatial and textual alignment of web content.
    • Distributions: The median document contains 1 image and 677 text tokens. The maximum number of images allowed per document is capped at 30, and documents containing zero images are excluded.
  2. Knowl 2 — OBELICS Multimodal Extraction and Filtering Pipeline

    algorithm

    The construction of the OBELICS corpus processes 41.2 billion raw web pages from 25 Common Crawl dumps through four successive transformation and filtering stages:

    1. Text Filtering and Quality Classification (41.2B →\rightarrow 1.1B documents):

      • Non-English text is removed using a FastText language classifier, filtering out 63.6% of raw documents.
      • Near-duplicate content is pruned via MinHash deduplication.
      • Documents with excessive n-gram or paragraph repetition (based on MassiveText heuristics) are eliminated.
      • High-quality human-written text is identified using a fast logistic regression classifier trained on hashed token frequencies, with Wikipedia and OpenWebText as positive examples and raw Common Crawl as negative examples.
    2. HTML DOM Tree Simplification (1.1B documents):

      • Raw HTML pages are parsed into Document Object Model (DOM) trees.
      • Heuristic rules unwrap semantic tags, eliminate boilerplate template structures, discard advertisements, navigation bars, and headers/footers, and strip out generic images (e.g., logos). This reduces the raw HTML size by over 10×10\times.
    3. Multimodal Document Extraction:

      • Text paragraphs and image URLs are extracted sequentially from the simplified DOM tree, preserving natural layout line breaks and visual flow.
      • From 3.6 billion identified image URLs, 55% (~2.0 billion images) are successfully downloaded.
    4. Granular Filtering (1.1B docs, 2.0B images →\rightarrow 365M docs, 1.4B images):

      • Node-level image filtering: Images with extreme aspect ratios, excessively small dimensions, excessively large dimensions, or URLs containing banned substrings (such as logo) are removed.
      • Paragraph-level text filtering: Paragraphs are discarded if they fail minimum word counts, exhibit high character/word repetition, have low punctuation or stop-word frequencies, possess high perplexity under a KenLM 5-gram language model trained on Wikipedia, contain flagged adult terms, or feature boilerplate social-sharing phrases.
      • Document-level filtering: Documents with 0 images or more than 30 images are eliminated, alongside documents failing aggregate text quality thresholds.
  3. Knowl 3 — Responsible Filtering and Multi-Level Deduplication in OBELICS

    model/method

    To transition from the filtered pool of 365 million documents and 1.4 billion images to the final OBELICS dataset, responsible content filters and hierarchical deduplication protocols are applied:

    • Creator Opt-Out Compliance: All image URLs are checked against the Spawning API to respect content creator preferences, removing all images whose rights holders opted out of AI training.
    • URL-Based Image Frequency Capping: Images appearing more than 10 times across the corpus (typically persistent web icons, trackers, and ubiquitous ad banners) are discarded, and duplicate images within the same document are removed.
    • Adult Content Removal: Documents containing any image flagged as pornographic by an open-source NSFW image classifier, or containing banned adult substrings in image URLs, are purged entirely.
    • Document Deduplication: Web documents are clustered by URL and by identical image sets, retaining only the most recent crawl instance in each cluster.
    • Intra-Domain Paragraph Deduplication: Exact-match paragraphs repeated across different web pages sharing the same domain name (such as cookie policies, legal disclaimers, or navigation footers) are eliminated, removing approximately 15% of all extracted text.
  4. Knowl 4 — Comparison of OBELICS with Existing Large-Scale Multimodal Corpora

    data/table

    The scale, image uniqueness, and openness of OBELICS are compared against the primary closed and open interleaved multimodal datasets:

    Dataset Images % Unique Images Documents Tokens Open
    KOSMOS-1 - - 71M - No
    M3W (Flamingo) 185M - 43M - No
    mmc4-ff 385M 60.6% 79M 34B Yes
    mmc4 585M - 103M 43B Yes
    OBELICS 353M 84.3% 141M 115B Yes

    OBELICS provides the largest number of documents (141M) and text tokens (115B) among open and closed interleaved datasets. While the full mmc4 corpus contains a higher gross count of images (585M), it exhibits significant duplication (only 60.6% unique images in mmc4-ff compared to 84.3% in OBELICS). Furthermore, mmc4 does not cap the maximum number of images per document, leading to uneven image distributions concentrated in low-text, image-heavy spam documents, whereas OBELICS enforces a 30-image maximum per document.

  5. Knowl 5 — Few-Shot Multimodal Benchmark Performance of IDEFICS vs Flamingo and OpenFlamingo

    data/table

    The evaluation compares IDEFICS-9B and IDEFICS-80B (trained on OBELICS, LAION, open captioning datasets, and Wikipedia) against Flamingo and OpenFlamingo (trained on mmc4) across 0, 4, 8, 16, and 32 in-context example shots in an open-ended setting with random demonstrations:

    Model Shot COCO Flickr30k VQAv2 OKVQA TextVQA VizWiz VisDial HatefulMemes
    Flamingo-9B 0 79.4 61.5 51.8 44.7 31.8 22.8 48.0 57.0
    OpenFlamingo-9B 0 79.5 59.5 52.7 37.8 24.2 27.5 - 51.6
    IDEFICS-9B 0 46.0 27.3 50.9 38.4 25.9 35.5 48.7 51.8
    Flamingo-9B 4 93.1 72.6 56.3 49.3 33.6 34.9 50.4 62.7
    OpenFlamingo-9B 4 89.0 65.8 54.8 40.1 28.2 34.1 - 54.0
    IDEFICS-9B 4 93.0 59.7 55.4 45.4 27.6 36.9 47.9 50.7
    Flamingo-9B 8 99.0 73.4 58.0 50.0 33.6 39.4 51.2 63.9
    OpenFlamingo-9B 8 96.3 62.9 54.8 41.1 29.1 38.5 - 54.7
    IDEFICS-9B 8 97.0 61.9 56.4 47.7 27.5 40.4 47.6 51.1
    Flamingo-9B 16 102.2 72.7 59.4 50.8 33.5 43.0 51.3 64.5
    OpenFlamingo-9B 16 98.8 62.8 54.3 42.7 27.3 42.5 - 53.9
    IDEFICS-9B 16 99.7 64.5 57.0 48.4 27.9 42.6 - 50.1
    Flamingo-9B 32 106.3 72.8 60.4 51.0 32.6 44.0 50.4 63.5
    OpenFlamingo-9B 32 99.5 61.3 53.3 42.4 23.8 44.0 - 53.8
    IDEFICS-9B 32 98.0 64.3 57.9 49.6 28.3 43.7 - 49.8
    Flamingo-80B 0 84.3 67.2 56.3 50.6 35.0 31.6 52.0 46.4
    IDEFICS-80B 0 91.8 53.7 60.0 45.2 30.9 36.0 48.9 60.6
    Flamingo-80B 4 103.2 75.1 63.1 57.4 36.5 39.6 55.6 68.6
    IDEFICS-80B 4 110.3 73.7 63.6 52.4 34.4 40.4 48.4 57.8
    Flamingo-80B 8 108.8 78.2 65.6 57.5 37.3 44.8 56.4 70.0
    IDEFICS-80B 8 114.3 76.6 64.8 55.1 35.7 46.1 47.9 58.2
    Flamingo-80B 16 110.5 78.9 66.8 57.8 37.6 48.4 56.8 70.0
    IDEFICS-80B 16 116.6 80.1 65.4 56.8 36.3 48.3 - 57.8
    Flamingo-80B 32 113.8 75.4 67.6 57.8 37.9 49.8 55.6 70.0
    IDEFICS-80B 32 116.6 81.1 65.9 57.8 36.7 50.0 - 52.5

    Evaluation tasks, metrics, and splits:

    • COCO: CIDEr, test split.
    • Flickr30k: CIDEr, Karpathy test split.
    • VQAv2: VQA accuracy, test-dev split.
    • OKVQA: VQA accuracy, val split.
    • TextVQA: VQA accuracy, val split.
    • VizWiz: VQA accuracy, test-dev split.
    • VisDial: NDCG, val split.
    • HatefulMemes: ROC-AUC, test seen split.

    IDEFICS-80B matches or exceeds Flamingo-80B on 4 out of 8 benchmarks at 32 shots (COCO: 116.6 vs 113.8; Flickr30k: 81.1 vs 75.4; OKVQA: 57.8 vs 57.8; VizWiz: 50.0 vs 49.8). At the 9B parameter scale, IDEFICS-9B achieves an aggregate score of 56.5 across tasks compared to 55.8 for OpenFlamingo-9B (taking the best shot count per task).

  6. Knowl 6 — Sample Efficiency of Interleaved Documents versus Image-Text Pairs

    empirical result

    Training vision-language models on interleaved multimodal web documents (OBELICS) provides substantial sample efficiency advantages over training on isolated image-text pairs (LAION):

    • Sample Efficiency: Models trained on interleaved web documents reach equivalent downstream benchmark performance using an order of magnitude fewer training images than models trained on image-text pairs, even though the visual data in both corpora originates from Common Crawl. This efficiency stems from the richer, longer textual context surrounding images in natural web pages.
    • Task Dynamics: On visual question answering (VQA) benchmarks, the performance of models trained solely on image-text pairs initially increases and subsequently degrades as training progresses, whereas models trained on interleaved documents exhibit sustained improvement.
    • Modality Strengths: Image-text pairs provide a minor advantage on isolated image captioning, classification, and OCR tasks (where captions function effectively as fuzzy class labels).
    • Mixture Benefits: Combining interleaved multimodal documents (OBELICS) with image-text pairs (LAION) achieves higher aggregate multimodal performance across all compute, token, and image budgets than training on either data format in isolation.
  7. Knowl 7 — Text Quality and Content Analysis of OBELICS

    empirical result

    The text quality, topic distribution, and visual characteristics of OBELICS were analyzed through perplexity evaluation, topic modeling, and qualitative auditing:

    • Perplexity Benchmark: Text perplexity was evaluated on 100,000 document samples using a 5-gram language model trained on Wikipedia. OBELICS exhibits significantly lower average perplexity than web corpora such as c4, mmc4, and OSCAR, matching the quality distribution of The Pile (a human-curated corpus).
    • Topic and Domain Distribution: Latent Dirichlet Allocation (LDA) across 20 and 200 topics reveals broad subject coverage spanning Politics, Health, Music, Science, and Technology, with news media websites representing the most frequent domain sources.
    • Sample Audit (250 Documents):
      • 80% of documents contain photographic images, and 29% contain graphical illustrations or cartoons.
      • 90% of documents exhibit clear alignment between all embedded images and the accompanying text.
      • 30% of documents contain images with embedded written words, and 5% feature structured document imagery (tables, slides, scanned documents), providing natural OCR pretraining signals.
      • 46% of documents depict human faces.
      • 0% contain NSFW or pornographic content, 3% contain watermarks, and 2% contain logos.
      • 7% contain residual uncaptured artifacts (such as minor non-English text fragments or spam).
  8. Knowl 8 — IDEFICS Multimodal Architecture and Training Framework

    model/method

    IDEFICS is an open multimodal vision-language model replicating the Flamingo architecture, available in 9-billion and 80-billion parameter configurations:

    • Backbone Architecture: Combines a frozen autoregressive language model backbone (LLaMA) and a frozen visual encoder backbone (OpenCLIP).
    • Cross-Modal Fusion: Variable-length spatial visual embeddings from OpenCLIP are condensed into a fixed set of visual tokens using a Transformer-based Perceiver Resampler. These visual tokens are then integrated into the language backbone via learnable gated cross-attention Transformer layers interspersed among the frozen language layers.
    • Sequence Construction: Interleaved web documents are converted directly into token sequences representing the ordered progression of text paragraphs and visual tokens. For image-text pair datasets, sequences are constructed by packing multiple images alongside their corresponding text captions.
    • Training Objective: Standard causal language modeling objective minimizing the cross-entropy loss for next-token prediction over text tokens, conditioned on all preceding textual and visual tokens.

Coverage note — Appendix details containing exhaustive lists of HTML DOM tag rules, individual banned keyword lists, and extended ablation hyperparameter grids were omitted as the primary contributions and results are fully captured by the knowls above.

References

  1. 1.Abbas, A., K. Tirumala, D. Simig, S. Ganguli, and A. S. Morcos (2023). Semdedup: Data-efficient learning at web-scale through semantic deduplication.
  2. 2.Aghajanyan, A., B. Huang, C. Ross, V. Karpukhin, H. Xu, N. Goyal, D. Okhonko, M. Joshi, G. Ghosh, M. Lewis, and L. Zettlemoyer (2022). Cm3: A causal masked multimodal model of the internet. ArXiv abs/2201.07520.
  3. 3.Alayrac, J.-B., J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. a. Bińkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan (2022). Flamingo: a visual language model for few-shot learning. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Advances in Neural Information Processing Systems, Volume 35, pp. 23716–23736. Curran Associates, Inc.
  4. 4.Antol, S., A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh (2015). VQA: Visual Question Answering. In International Conference on Computer Vision (ICCV).
  5. 5.Awadalla, A., I. Gao, J. Gardner, J. Hessel, Y. Hanafy, W. Zhu, K. Marathe, Y. Bitton, S. Gadre, S. Sagawa, J. Jitsev, S. Kornblith, P. W. Koh, G. Ilharco, M. Wortsman, and L. Schmidt (2023). Openflamingo: An open-source framework for training large autoregressive vision-language models.
  6. 6.Bai, Y., A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan (2022). Training a helpful and harmless assistant with reinforcement learning from human feedback.
  7. 7.Beaumont, R. (2021). img2dataset: Easily turn large sets of image urls to an image dataset. https://github.com/rom1504/img2dataset.
  8. 8.Bender, E. M., T. Gebru, A. McMillan-Major, and S. Shmitchell (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency.
  9. 9.Biderman, S. and W. J. Scheirer (2020, 12 Dec). Pitfalls in machine learning research: Reexamining the development cycle. In J. Zosa Forde, F. Ruiz, M. F. Pradier, and A. Schein (Eds.), Proceedings on "I Can't Believe It's Not Better!" at NeurIPS Workshops, Volume 137 of Proceedings of Machine Learning Research, pp. 106–117. PMLR.
  10. 10.Biderman, S., H. Schoelkopf, Q. Anthony, H. Bradley, K. O'Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, A. Skowron, L. Sutawika, and O. van der Wal (2023). Pythia: A suite for analyzing large language models across training and scaling.
  11. 11.Birhane, A., V. U. Prabhu, and E. Kahembwe (2021). Multimodal datasets: misogyny, pornography, and malignant stereotypes. ArXiv abs/2110.01963.
  12. 12.Blei, D. M., A. Y. Ng, and M. I. Jordan (2003, mar). Latent dirichlet allocation. J. Mach. Learn. Res. 3 (null), 993–1022.
  13. 13.Broder, A. (1997). On the resemblance and containment of documents. In Proceedings. Compression and Complexity of SEQUENCES 1997 (Cat. No.97TB100171), pp. 21–29.
  14. 14.Brown, T., B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei (2020). Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin (Eds.), Advances in Neural Information Processing Systems, Volume 33, pp. 1877–1901. Curran Associates, Inc.
  15. 15.Byeon, M., B. Park, H. Kim, S. Lee, W. Baek, and S. Kim (2022). Coyo-700m: Image-text pair dataset. https://github.com/kakaobrain/coyo-dataset.
  16. 16.Caswell, I., T. Breiner, D. van Esch, and A. Bapna (2020). Language id in the wild: Unexpected challenges on the path to a thousand-language web text corpus. ArXiv abs/2010.14571.
  17. 17.Changpinyo, S., P. Sharma, N. Ding, and R. Soricut (2021). Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR.
  18. 18.Chowdhery, A., S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel (2022). Palm: Scaling language modeling with pathways.
  19. 19.Das, A., S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. F. Moura, D. Parikh, and D. Batra (2017, July). Visual dialog. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  20. 20.Dehghani, M., J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, R. Jenatton, L. Beyer, M. Tschannen, A. Arnab, X. Wang, C. Riquelme, M. Minderer, J. Puigcerver, U. Evci, M. Kumar, S. van Steenkiste, G. F. Elsayed, A. Mahendran, F. Yu, A. Oliver, F. Huot, J. Bastings, M. P. Collier, A. Gritsenko, V. Birodkar, C. Vasconcelos, Y. Tay, T. Mensink, A. Kolesnikov, F. Pavetić, D. Tran, T. Kipf, M. Lučić, X. Zhai, D. Keysers, J. Harmsen, and N. Houlsby (2023). Scaling vision transformers to 22 billion parameters.
  21. 21.Deng, X., P. Shiralkar, C. Lockard, B. Huang, and H. Sun (2022). Dom-lm: Learning generalizable representations for html documents. ArXiv abs/2201.10608.
  22. 22.Desai, K., G. Kaul, Z. Aysola, and J. Johnson (2021). Redcaps: Web-curated image-text data created by the people, for the people. In J. Vanschoren and S. Yeung (Eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, Volume 1. Curran.
  23. 23.Dodge, J., A. Marasović, G. Ilharco, D. Groeneveld, M. Mitchell, and M. Gardner (2021). Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Conference on Empirical Methods in Natural Language Processing.
  24. 24.Gadre, S. Y., G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, E. Orgad, R. Entezari, G. Daras, S. Pratt, V. Ramanujan, Y. Bitton, K. Marathe, S. Mussmann, R. Vencu, M. Cherti, R. Krishna, P. W. Koh, O. Saukh, A. Ratner, S. Song, H. Hajishirzi, A. Farhadi, R. Beaumont, S. Oh, A. Dimakis, J. Jitsev, Y. Carmon, V. Shankar, and L. Schmidt (2023). Datacomp: In search of the next generation of multimodal datasets. arXiv preprint arXiv:2304.14108.
  25. 25.Gao, L., S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy (2020). The Pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  26. 26.Gokaslan, A. and V. Cohen (2019). Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus.
  27. 27.Gu, J., X. Meng, G. Lu, L. Hou, N. Minzhe, X. Liang, L. Yao, R. Huang, W. Zhang, X. Jiang, C. XU, and H. Xu (2022). Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Advances in Neural Information Processing Systems, Volume 35, pp. 26418–26431. Curran Associates, Inc.
  28. 28.Gurari, D., Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham (2018). Vizwiz grand challenge: Answering visual questions from blind people.
  29. 29.Heafield, K. (2011, July). KenLM: Faster and smaller language model queries. In Proceedings of the Sixth Workshop on Statistical Machine Translation, Edinburgh, Scotland, pp. 187–197. Association for Computational Linguistics.
  30. 30.Hoffmann, J., S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre (2022). Training compute-optimal large language models.
  31. 31.Huang, S., L. Dong, W. Wang, Y. Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, B. Patra, Q. Liu, K. Aggarwal, Z. Chi, J. Bjorck, V. Chaudhary, S. Som, X. Song, and F. Wei (2023). Language is not all you need: Aligning perception with language models.
  32. 32.Jaegle, A., F. Gimeno, A. Brock, A. Zisserman, O. Vinyals, and J. Carreira (2021). Perceiver: General perception with iterative attention.
  33. 33.Jia, C., Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. V. Le, Y.-H. Sung, Z. Li, and T. Duerig (2021). Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning.
  34. 34.Jiang, A. Q., S. Welleck, J. P. Zhou, T. Lacroix, J. Liu, W. Li, M. Jamnik, G. Lample, and Y. Wu (2023). Draft, sketch, and prove: Guiding formal theorem provers with informal proofs. In The Eleventh International Conference on Learning Representations.
  35. 35.Joulin, A., E. Grave, P. Bojanowski, and T. Mikolov (2017, April). Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, Valencia, Spain, pp. 427–431. Association for Computational Linguistics.
  36. 36.Kaplan, J., S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei (2020). Scaling laws for neural language models.
  37. 37.Kärkkäinen, K. and J. Joo (2021). Fairface: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation. 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), 1547–1557.
  38. 38.Kiela, D., H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, and D. Testuggine (2020). The hateful memes challenge: Detecting hate speech in multimodal memes. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin (Eds.), Advances in Neural Information Processing Systems, Volume 33, pp. 2611–2624. Curran Associates, Inc.
  39. 39.Koh, J. Y., R. Salakhutdinov, and D. Fried (2023). Grounding language models to images for multimodal generation.
  40. 40.Laborde, G. Deep nn for nsfw detection.
  41. 41.Laurençon, H., L. Saulnier, T. Wang, C. Akiki, A. Villanova del Moral, T. Le Scao, L. Von Werra, C. Mou, E. González Ponferrada, H. Nguyen, J. Frohberg, M. Šaško, Q. Lhoest, A. McMillan-Major, G. Dupont, S. Biderman, A. Rogers, L. Ben allal, F. De Toni, G. Pistilli, O. Nguyen, S. Nikpoor, M. Masoud, P. Colombo, J. de la Rosa, P. Villegas, T. Thrush, S. Longpre, S. Nagel, L. Weber, M. Muñoz, J. Zhu, D. Van Strien, Z. Alyafeai, K. Almubarak, M. C. Vu, I. Gonzalez-Dios, A. Soroa, K. Lo, M. Dey, P. Ortiz Suarez, A. Gokaslan, S. Bose, D. Adelani, L. Phan, H. Tran, I. Yu, S. Pai, J. Chim, V. Lepercq, S. Ilic, M. Mitchell, S. A. Luccioni, and Y. Jernite (2022). The bigscience roots corpus: A 1.6tb composite multilingual dataset. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Advances in Neural Information Processing Systems, Volume 35, pp. 31809–31826. Curran Associates, Inc.
  42. 42.Lee, K., D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini (2022). Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics.
  43. 43.Li, J., D. Li, S. Savarese, and S. Hoi (2023). Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.
  44. 44.Li, J., D. Li, C. Xiong, and S. Hoi (2022). Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML.
  45. 45.Li, R., L. B. Allal, Y. Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim, Q. Liu, E. Zheltonozhskii, T. Y. Zhuo, T. Wang, O. Dehaene, M. Davaadorj, J. Lamy-Poirier, J. Monteiro, O. Shliazhko, N. Gontier, N. Meade, A. Zebaze, M.-H. Yee, L. K. Umapathi, J. Zhu, B. Lipkin, M. Oblokulov, Z. Wang, R. Murthy, J. Stillerman, S. S. Patel, D. Abulkhanov, M. Zocca, M. Dey, Z. Zhang, N. Fahmy, U. Bhattacharyya, W. Yu, S. Singh, S. Luccioni, P. Villegas, M. Kunakov, F. Zhdanov, M. Romero, T. Lee, N. Timor, J. Ding, C. Schlesinger, H. Schoelkopf, J. Ebert, T. Dao, M. Mishra, A. Gu, J. Robinson, C. J. Anderson, B. Dolan-Gavitt, D. Contractor, S. Reddy, D. Fried, D. Bahdanau, Y. Jernite, C. M. Ferrandis, S. Hughes, T. Wolf, A. Guha, L. von Werra, and H. de Vries (2023). Starcoder: may the source be with you!
  46. 46.Lin, T.-Y., M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C. L. Zitnick, and P. Dollár (2014). Microsoft coco: Common objects in context. cite arxiv:1405.0312Comment: 1) updated annotation pipeline description and figures; 2) added new section describing datasets splits; 3) updated author list.
  47. 47.Liu, S., L. Fan, E. Johns, Z. Yu, C. Xiao, and A. Anandkumar (2023). Prismer: A vision-language model with an ensemble of experts. arXiv preprint arXiv:2303.02506.
  48. 48.Liu, Y., G. Zhu, B. Zhu, Q. Song, G. Ge, H. Chen, G. Qiao, R. Peng, L. Wu, and J. Wang (2022). Taisu: A 166m large-scale high-quality dataset for chinese vision-language pre-training. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Advances in Neural Information Processing Systems, Volume 35, pp. 16705–16717. Curran Associates, Inc.
  49. 49.Loshchilov, I. and F. Hutter (2017). Fixing weight decay regularization in adam. CoRR abs/1711.05101.
  50. 50.Luccioni, A. S., C. Akiki, M. Mitchell, and Y. Jernite (2023). Stable bias: Analyzing societal representations in diffusion models.
  51. 51.Marino, K., M. Rastegari, A. Farhadi, and R. Mottaghi (2019). Ok-vqa: A visual question answering benchmark requiring external knowledge. In Conference on Computer Vision and Pattern Recognition (CVPR).
  52. 52.Mishra, A., K. Alahari, and C. V. Jawahar (2012). Scene text recognition using higher order language priors. In BMVC.
  53. 53.Nichol, A., P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen (2022). Glide: Towards photorealistic image generation and editing with text-guided diffusion models.
  54. 54.Ortiz Suárez, P. J., L. Romary, and B. Sagot (2020, July). A monolingual approach to contextualized word embeddings for mid-resource languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, pp. 1703–1714. Association for Computational Linguistics.
  55. 55.Ouyang, L., J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe (2022). Training language models to follow instructions with human feedback. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Advances in Neural Information Processing Systems, Volume 35, pp. 27730–27744. Curran Associates, Inc.
  56. 56.Piktus, A., C. Akiki, P. Villegas, H. Laurençon, G. Dupont, A. S. Luccioni, Y. Jernite, and A. Rogers (2023). The roots search tool: Data transparency for llms.
  57. 57.Radenovic, F., A. Dubey, A. Kadian, T. Mihaylov, S. Vandenhende, Y. Patel, Y. Wen, V. Ramanathan, and D. Mahajan (2023). Filtering, distillation, and hard negatives for vision-language pre-training.
  58. 58.Radford, A., J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever (2021). Learning transferable visual models from natural language supervision. In International Conference on Machine Learning.
  59. 59.Rae, J. W., S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, P.-S. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A. Wu, E. Elsen, S. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J.-B. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d'Autume, Y. Li, T. Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, B. Hechtman, L. Weidinger, I. Gabriel, W. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving (2022). Scaling language models: Methods, analysis & insights from training gopher.
  60. 60.Raffel, C., N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu (2019). Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv e-prints.
  61. 61.Ramesh, A., P. Dhariwal, A. Nichol, C. Chu, and M. Chen (2022). Hierarchical text-conditional image generation with clip latents.
  62. 62.Rombach, R., A. Blattmann, D. Lorenz, P. Esser, and B. Ommer (2021). High-resolution image synthesis with latent diffusion models.
  63. 63.Saharia, C., W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, S. K. S. Ghasemipour, B. K. Ayan, S. S. Mahdavi, R. G. Lopes, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi (2022). Photorealistic text-to-image diffusion models with deep language understanding.
  64. 64.Schuhmann, C., R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev (2022). Laion-5b: An open large-scale dataset for training next generation image-text models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Advances in Neural Information Processing Systems, Volume 35, pp. 25278–25294. Curran Associates, Inc.
  65. 65.Schuhmann, C., R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki (2021). Laion-400m: Open dataset of clip-filtered 400 million image-text pairs.
  66. 66.Sharma, P., N. Ding, S. Goodman, and R. Soricut (2018). Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of ACL.
  67. 67.Singh, A., R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela (2022). FLAVA: A foundational language and vision alignment model. In CVPR.
  68. 68.Singh, A., V. Natarjan, M. Shah, Y. Jiang, X. Chen, D. Parikh, and M. Rohrbach (2019). Towards vqa models that can read. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8317–8326.
  69. 69.Sorscher, B., R. Geirhos, S. Shekhar, S. Ganguli, and A. Morcos (2022). Beyond neural scaling laws: beating power law scaling via data pruning. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Advances in Neural Information Processing Systems, Volume 35, pp. 19523–19536. Curran Associates, Inc.
  70. 70.Srinivasan, K., K. Raman, J. Chen, M. Bendersky, and M. Najork (2021). Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '21, New York, NY, USA, pp. 2443–2449. Association for Computing Machinery.
  71. 71.Team, M. N. (2023). Introducing mpt-7b: A new standard for open-source, commercially usable llms.
  72. 72.Touvron, H., T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample (2023). Llama: Open and efficient foundation language models.
  73. 73.Wang, P., A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang (2022, 17–23 Jul). OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato (Eds.), Proceedings of the 39th International Conference on Machine Learning, Volume 162 of Proceedings of Machine Learning Research, pp. 23318–23340. PMLR.
  74. 74.Wang, Q., Y. Fang, A. Ravula, F. Feng, X. Quan, and D. Liu (2022). Webformer: The web-page transformer for structure information extraction.
  75. 75.Wang, W., H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som, and F. Wei (2022). Image as a foreign language: Beit pretraining for all vision and vision-language tasks.
  76. 76.Webster, R., J. Rabin, L. Simon, and F. Jurie (2023). On the de-duplication of laion-2b.
  77. 77.Workshop, B., :, T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilić, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, M. Gallé, J. Tow, A. M. Rush, S. Biderman, A. Webson, P. S. Ammanamanchi, T. Wang, B. Sagot, N. Muennighoff, A. V. del Moral, O. Ruwase, R. Bawden, S. Bekman, A. McMillan-Major, I. Beltagy, H. Nguyen, L. Saulnier, S. Tan, P. O. Suarez, V. Sanh, H. Laurençon, Y. Jernite, J. Launay, M. Mitchell, C. Raffel, A. Gokaslan, A. Simhi, A. Soroa, A. F. Aji, A. Alfassy, A. Rogers, A. K. Nitzav, C. Xu, C. Mou, C. Emezue, C. Klamm, C. Leong, D. van Strien, D. I. Adelani, D. Radev, E. G. Ponferrada, E. Levkovizh, E. Kim, E. B. Natan, F. D. Toni, G. Dupont, G. Kruszewski, G. Pistilli, H. Elsahar, H. Benyamina, H. Tran, I. Yu, I. Abdulmumin, I. Johnson, I. Gonzalez-Dios, J. de la Rosa, J. Chim, J. Dodge, J. Zhu, J. Chang, J. Frohberg, J. Tobing, J. Bhattacharjee, K. Almubarak, K. Chen, K. Lo, L. V. Werra, L. Weber, L. Phan, L. B. allal, L. Tanguy, M. Dey, M. R. Muñoz, M. Masoud, M. Grandury, M. Šaško, M. Huang, M. Coavoux, M. Singh, M. T.-J. Jiang, M. C. Vu, M. A. Jauhar, M. Ghaleb, N. Subramani, N. Kassner, N. Khamis, O. Nguyen, O. Espejel, O. de Gibert, P. Villegas, P. Henderson, P. Colombo, P. Amuok, Q. Lhoest, R. Harliman, R. Bommasani, R. L. López, R. Ribeiro, S. Osei, S. Pyysalo, S. Nagel, S. Bose, S. H. Muhammad, S. Sharma, S. Longpre, S. Nikpoor, S. Silberberg, S. Pai, S. Zink, T. T. Torrent, T. Schick, T. Thrush, V. Danchev, V. Nikoulina, V. Laippala, V. Lepercq, V. Prabhu, Z. Alyafeai, Z. Talat, A. Raja, B. Heinzerling, C. Si, D. E. Tañar, E. Salesky, S. J. Mielke, W. Y. Lee, A. Sharma, A. Santilli, A. Chaffin, A. Stiegler, D. Datta, E. Szczechla, G. Chhablani, H. Wang, H. Pandey, H. Strobelt, J. A. Fries, J. Rozen, L. Gao, L. Sutawika, M. S. Bari, M. S. Al-shaibani, M. Manica, N. Nayak, R. Teehan, S. Albanie, S. Shen, S. Ben-David, S. H. Bach, T. Kim, T. Bers, T. Fevry, T. Neeraj, U. Thakker, V. Raunak, X. Tang, Z.-X. Yong, Z. Sun, S. Brody, Y. Uri, H. Tojarieh, A. Roberts, H. W. Chung, J. Tae, J. Phang, O. Press, C. Li, D. Narayanan, H. Bourfoune, J. Casper, J. Rasley, M. Ryabinin, M. Mishra, M. Zhang, M. Shoeybi, M. Peyrounette, N. Patry, N. Tazi, O. Sanseviero, P. von Platen, P. Cornette, P. F. Lavallée, R. Lacroix, S. Rajbhandari, S. Gandhi, S. Smith, S. Requena, S. Patil, T. Dettmers, A. Baruwa, A. Singh, A. Cheveleva, A.-L. Ligozat, A. Subramonian, A. Névéol, C. Lovering, D. Garrette, D. Tunuguntla, E. Reiter, E. Taktasheva, E. Voloshina, E. Bogdanov, G. I. Winata, H. Schoelkopf, J.-C. Kalo, J. Novikova, J. Z. Forde, J. Clive, J. Kasai, K. Kawamura, L. Hazan, M. Carpuat, M. Clinciu, N. Kim, N. Cheng, O. Serikov, O. Antverg, O. van der Wal, R. Zhang, R. Zhang, S. Gehrmann, S. Mirkin, S. Pais, T. Shavrina, T. Scialom, T. Yun, T. Limisiewicz, V. Rieser, V. Protasov, V. Mikhailov, Y. Pruksachatkun, Y. Belinkov, Z. Bamberger, Z. Kasner, A. Rueda, A. Pestana, A. Feizpour, A. Khan, A. Faranak, A. Santos, A. Hevia, A. Unldreaj, A. Aghagol, A. Abdollahi, A. Tammour, A. HajiHosseini, B. Behroozi, B. Ajibade, B. Saxena, C. M. Ferrandis, D. Contractor, D. Lansky, D. David, D. Kiela, D. A. Nguyen, E. Tan, E. Baylor, E. Ozoani, F. Mirza, F. Ononiwu, H. Rezanejad, H. Jones, I. Bhattacharya, I. Solaiman, I. Sedenko, I. Nejadgholi, J. Passmore, J. Seltzer, J. B. Sanz, L. Dutra, M. Samagaio, M. Elbadri, M. Mieskes, M. Gerchick, M. Akinlolu, M. McKenna, M. Qiu, M. Ghauri, M. Burynok, N. Abrar, N. Rajani, N. Elkott, N. Fahmy, O. Samuel, R. An, R. Kromann, R. Hao, S. Alizadeh, S. Shubber, S. Wang, S. Roy, S. Viguier, T. Le, T. Oyebade, T. Le, Y. Yang, Z. Nguyen, A. R. Kashyap, A. Palasciano, A. Callahan, A. Shukla, A. Miranda-Escalada, A. Singh, B. Beilharz, B. Wang, C. Brito, C. Zhou, C. Jain, C. Xu, C. Fourrier, D. L. Periñán, D. Molano, D. Yu, E. Manjavacas, F. Barth, F. Fuhrimann, G. Altay, G. Bayrak, G. Burns, H. U. Vrabec, I. Bello, I. Dash, J. Kang, J. Giorgi, J. Golde, J. D. Posada, K. R. Sivaraman, L. Bulchandani, L. Liu, L. Shinzato, M. H. de Bykhovetz, M. Takeuchi, M. Pàmies, M. A. Castillo, M. Nezhurina, M. Sänger, M. Samwald, M. Cullan, M. Weinberg, M. D. Wolf, M. Mihaljcic, M. Liu, M. Freidank, M. Kang, N. Seelam, N. Dahlberg, N. M. Broad, N. Muellner, P. Fung, P. Haller, R. Chandrasekhar, R. Eisenberg, R. Martin, R. Canalli, R. Su, R. Su, S. Cahyawijaya, S. Garda, S. S. Deshmukh, S. Mishra, S. Kiblawi, S. Ott, S. Sang-aroonsiri, S. Kumar, S. Schweter, S. Bharati, T. Laud, T. Gigant, T. Kainuma, W. Kusa, Y. Labrak, Y. S. Bajaj, Y. Venkatraman, Y. Xu, Y. Xu, Y. Xu, Z. Tan, Z. Xie, Z. Ye, M. Bras, Y. Belkada, and T. Wolf (2023). Bloom: A 176b-parameter open-access multilingual language model.
  78. 78.Xie, S. M., H. Pham, X. Dong, N. Du, H. Liu, Y. Lu, P. Liang, Q. V. Le, T. Ma, and A. W. Yu (2023). Doremi: Optimizing data mixtures speeds up language model pretraining.
  79. 79.Yang, Z., Z. Gan, J. Wang, X. Hu, Y. Lu, Z. Liu, and L. Wang (2022). An empirical study of gpt-3 for few-shot knowledge-based vqa. In Proceedings of the AAAI Conference on Artificial Intelligence, Volume 36, pp. 3081–3089.
  80. 80.Young, P., A. Lai, M. Hodosh, and J. Hockenmaier (2014). From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics 2, 67–78.
  81. 81.Yu, J., Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu (2022). Coca: Contrastive captioners are image-text foundation models. Transactions on Machine Learning Research.
  82. 82.Yuan, S., S. Zhao, J. Leng, Z. Xue, H. Zhao, P. Liu, Z. Gong, W. X. Zhao, J. Li, and J. Tang (2022). Wudaomm: A large-scale multi-modal dataset for pre-training models.
  83. 83.Yuksekgonul, M., F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou (2023). When and why vision-language models behave like bags-of-words, and what to do about it? In International Conference on Learning Representations.
  84. 84.Zhang, B. and R. Sennrich (2019). Root Mean Square Layer Normalization. In Advances in Neural Information Processing Systems 32, Vancouver, Canada.
  85. 85.Zhang, J., Y. Zhao, M. Saleh, and P. J. Liu (2019). Pegasus: Pre-training with extracted gap-sentences for abstractive summarization.
  86. 86.Zhang, R., J. Han, A. Zhou, X. Hu, S. Yan, P. Lu, H. Li, P. Gao, and Y. Qiao (2023). Llama-adapter: Efficient fine-tuning of language models with zero-init attention.
  87. 87.Zhang, S., S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer (2022). Opt: Open pre-trained transformer language models.
  88. 88.Zhou, Y., Y. Sheng, N. H. Vo, N. Edmonds, and S. Tata (2021). Simplified dom trees for transferable attribute extraction from the web. ArXiv abs/2101.02415.
  89. 89.Zhu, W., J. Hessel, A. Awadalla, S. Y. Gadre, J. Dodge, A. Fang, Y. Yu, L. Schmidt, W. Y. Wang, and Y. Choi (2023). Multimodal C4: An open, billion-scale corpus of images interleaved with text. arXiv preprint arXiv:2304.06939.

Citation

MLA
Laurençon, H., et al. “OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents”. arXiv, 2023, http://arxiv.org/abs/2306.16527v2.
APA
Laurençon, H., Saulnier, L., Tronchon, L., Bekman, S., Singh, A., Lozhkov, A., Wang, T., Karamcheti, S., Rush, A. M., Kiela, D., Cord, M., & Sanh, V. (2023). OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents. arXiv. http://arxiv.org/abs/2306.16527v2
Chicago
Laurençon, H., L. Saulnier, L. Tronchon, et al. 2023. “OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents”. arXiv. http://arxiv.org/abs/2306.16527v2.
Harvard
Laurençon, H. et al. (2023) “OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.16527v2.
Vancouver
1. Laurençon H, Saulnier L, Tronchon L, et al (2023) OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents. arXiv

BibTeX

@article{laurencon2023obelics,
  title = {OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents},
  author = {Laurençon, Hugo and Saulnier, Lucile and Tronchon, Léo and Bekman, Stas and Singh, Amanpreet and Lozhkov, Anton and Wang, Thomas and Karamcheti, Siddharth and Rush, Alexander M. and Kiela, Douwe and Cord, Matthieu and Sanh, Victor},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.16527v2},
  eprint = {2306.16527}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors