HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Antoine MiechDimitri ZhukovJean-Baptiste AlayracMakarand TapaswiIvan LaptevJosef Sivic
Introduces HowTo100M, a massive dataset of 136 million narrated video clips, demonstrating that models trained on automatically transcribed speech achieve superior performance on text-video retrieval and action localization across diverse domains without requiring manual captioning.
Developing artificial intelligence capable of understanding and connecting video with natural language—such as searching video archives by text description or localizing specific actions—normally requires enormous datasets of manually captioned video clips. Creating these datasets through human annotation is prohibitively expensive, time-consuming, subjective, and difficult to scale, which severely limits the scope and performance of existing models.
The article demonstrates that highly capable joint video-text representations can be trained entirely without manual annotations by using massive amounts of readily available instructional web videos paired with automatic speech transcriptions.
To achieve this, the authors collected HowTo100M, a dataset comprising 136 million video clips from 1.22 million narrated YouTube videos spanning over 23,000 physical tasks across 12 distinct categories. Using pre-extracted visual features and word embeddings, they trained a joint embedding model designed to map video clips and text into a shared semantic space. Crucially, the training process used an intra-video negative sampling strategy, pairing clips with incorrect narrations from the exact same video to force the model to focus on subtle action details rather than generic background cues.
The core findings show substantial performance and efficiency gains. First, the off-the-shelf model trained on HowTo100M establishes new state-of-the-art benchmarks on instructional video tasks: it achieved a 33.6% average recall for action step localization on the CrossTask benchmark (surpassing fully supervised baselines) and delivered top retrieval performance on cooking videos (YouCook2). Second, the model transfers remarkably well to non-instructional domains, such as generic web clips (MSR-VTT) and movie clips (LSMDC). Third, fine-tuning the pre-trained model on just 20% of the MSR-VTT dataset matched prior state-of-the-art performance trained on 100% of the data. Finally, empirical evaluations showed that model performance steadily improved as data volume grew, with no observed saturation at scale.
These results demonstrate that web-scale weakly supervised pre-training offers an effective, low-cost alternative to manual labeling, reducing the human annotation required to build strong domain-specific video models by up to 80%. Pre-training on large-scale instructional video consistently yields better downstream models across diverse video domains compared to training from scratch or pre-training on smaller datasets.
Organizations developing video search, retrieval, or activity analysis tools should adopt large-scale narrated video pre-training as a foundational base before fine-tuning on specialized downstream tasks. Rather than investing heavily in extensive manual captioning programs, teams can maximize efficiency by curating small, high-quality target datasets for domain-specific fine-tuning.
These conclusions come with certain caveats. Automatically transcribed narrations are inherently noisy; manual inspection showed that only about 51% of narrated clips strictly depict the spoken object or action. Additionally, because the source data focuses on physical how-to tasks, zero-shot performance on drastically different video formats, such as cinema, remains limited without fine-tuning. Despite this noise, the authors express high confidence that massive scale compensates for imperfect alignment, establishing a highly reliable and cost-effective methodology for video-language modeling.
- Paper: VideoBERT: A Joint Model for Video and Language Representation Learning, Chen Sun et al. (2019). VideoBERT establishes the foundational paradigm of learning multimodal video-language representations directly from spoken narrations via automated speech recognition on YouTube instructional videos.
- Paper: MSR-VTT: A Large Video Description Dataset for Bridging Video and Language, Jun Xu et al. (2016). MSR-VTT provides the primary open-domain benchmark that HowTo100M uses to evaluate the downstream transferability and retrieval performance of its learned embeddings.
- Paper: Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset, João Carreira et al. (2017). This paper introduces large-scale spatiotemporal video pre-training and standard 3D convolutional video backbones that underpin modern visual feature extraction in video-language models.
- Paper: Dense-Captioning Events in Videos, Ranjay Krishna et al. (2017). ActivityNet Captions formulated the benchmark and methodology for dense temporal event localization and descriptive captioning in untrimmed instructional and everyday videos.
- Paper: Learning Spatiotemporal Features with 3D Convolutional Networks, Du Tran et al. (2015). This seminal work establishes 3D convolutional networks as an effective method for extracting continuous spatiotemporal representations from video streams.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). Frozen in Time builds upon noisy video-text pretraining datasets like HowTo100M by designing an end-to-end visual transformer that jointly learns from both image and video captions.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). Video-LLaVA extends large-scale video-text alignment into instruction-tuned multimodal models by unifying image and video representations before projecting them to language backbones.
- Paper: Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, Hang Zhang et al. (2023). Video-LLaMA builds upon web-scale video-text retrieval representations to enable conversational large language models with joint audio and video understanding.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). ViViT advances video representation learning by replacing traditional 3D CNN feature extractors with pure spatial-temporal transformer architectures.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). Video-MME advances beyond basic video retrieval benchmarks by providing a comprehensive multi-modal evaluation suite for video understanding across complex domains and long durations.
