OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
Hugo LaurençonLucile SaulnierLéo TronchonStas BekmanAmanpreet SinghAnton LozhkovThomas WangSiddharth KaramchetiAlexander M. RushDouwe Kiela
Introduces the first fully open, web-scale dataset of interleaved image-text documents alongside competitive 9B and 80B multimodal models, enabling reproducible research and training for next-generation vision-language systems.
Modern vision and language artificial intelligence models achieve superior capabilities when trained on naturally structured web documents that interleave text and images, rather than on simple image-caption pairs. However, the large-scale multimodal web datasets previously used to train leading proprietary systems have remained private, limiting research replication, transparency, and broader development in the artificial intelligence community.
The article demonstrates the viability of creating an open, high-quality, web-scale multimodal document dataset, named OBELICS, and evaluates its effectiveness by training open vision and language models named IDEFICS at both 9-billion and 80-billion parameter scales.
The authors constructed the dataset from 25 raw web dumps spanning February 2020 to February 2023, originally containing 41.2 billion documents. The pipeline filtered out non-English content, stripped boilerplate web elements using webpage document structure rules, removed duplicate text and images, filtered out adult content, and respected creator opt-out preferences. The final dataset yielded 141 million multimodal documents, 353 million images, and 115 billion text tokens. To validate the resource, the authors trained multimodal models using this data combined with open image-text pairs and benchmarked their performance across eight standard evaluation tasks.
The analysis reveals several critical findings. First, training on interleaved web documents achieves equivalent multimodal performance using an order of magnitude fewer images compared to training solely on isolated image-text pairs. Second, the 80-billion parameter IDEFICS model matches or exceeds the performance of comparable leading proprietary models across multiple visual reasoning benchmarks. Third, at the 9-billion parameter scale, the model outperforms competing open-source alternatives trained on older, less filtered web corpuses. Finally, the textual content in OBELICS exhibits significantly higher quality and lower perplexity scores than existing open multimodal and general web baselines.
These findings indicate that organizations can train highly competitive multimodal systems using curated, openly accessible web documents without relying on proprietary training data. Leveraging structured, interleaved documents also lowers the data volume requirements for visual pre-training, reducing computational overhead and infrastructure costs while preserving document context.
Teams developing multimodal models should adopt curated interleaved document corpuses, ideally combining them with image-text pairs to balance visual question answering and fine-grained classification capabilities. Before deploying models broadly, practitioners should implement domain-specific guardrails, as residual web crawling risks like occasional uncaptured advertisements or public imagery remain present. Researchers should continue developing finer semantic filtering and multi-language extensions to expand dataset utility.
- Paper: LAION-5B: An open large-scale dataset for training next generation image-text models, Christoph Schuhmann et al. (2022). It provides the foundational methodology for large-scale open web scraping and multimodal filtering that OBELICS expands upon by moving from isolated image-text pairs to interleaved web documents.
- Paper: LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs, Christoph Schuhmann et al. (2021). It establishes the core pipeline for extracting and filtering hundreds of millions of multimodal pairs from Common Crawl, representing the preceding generation of open web-scale dataset curation.
- Paper: BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, Junnan Li et al. (2022). It introduces essential bootstrapping and data-filtering principles for vision-language pre-training that inform the quality filtering techniques applied to OBELICS.
- Paper: Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning, Piyush Sharma et al. (2018). It pioneers the web-scale extraction and heuristic filtering of raw HTML and alt-text into learnable vision-language pairs.
- Paper: Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts, Soravit Changpinyo et al. (2021). It examines how relaxing strict filtering rules on web-scraped image-text data improves open-vocabulary and long-tail visual recognition.
- Paper: Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision, Chao Jia et al. (2021). It demonstrates how massive, weakly filtered web corpora can effectively train large-scale vision-language models.
- Paper: The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only, Guilherme Penedo et al. (2023). It outlines document-structure heuristics and deduplication strategies for web-scale document filtering that directly parallel the pipeline design used in OBELICS.
- Paper: Generative Multimodal Models are In-Context Learners, Quan Sun et al. (2024). It leverages interleaved multimodal web pre-training data to scale up generative in-context learning across vision and language.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). It applies unified visual representations and curated data mixtures to transfer multimodal capabilities seamlessly across single-image, multi-image, and video domains.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). It extends open-source multimodal pre-training recipes by investigating joint linguistic-visual training paradigms and native parameter optimization.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). It builds on open multimodal foundation models by systematically scaling vision encoders, data curation recipes, and test-time reasoning.
- Paper: Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models, Siddharth Karamcheti et al. (2024). It empirically dissects the architectural and training design spaces of visually conditioned language models to optimize pre-training efficiency.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). It broadens the systematic evaluation of web data extraction and filtering pipelines established in multimodal datasets to language model pre-training testbeds.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). It scales multimodal pre-training architectures and long-context interleaved visual-textual sequences to advanced mixture-of-experts models.
