HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Antoine MiechDimitri ZhukovJean-Baptiste AlayracMakarand TapaswiIvan LaptevJosef Sivic

article2019ICCV1,519 citations

Introduces HowTo100M, a massive dataset of 136 million narrated video clips, demonstrating that models trained on automatically transcribed speech achieve superior performance on text-video retrieval and action localization across diverse domains without requiring manual captioning.

Listen

Developing artificial intelligence capable of understanding and connecting video with natural language—such as searching video archives by text description or localizing specific actions—normally requires enormous datasets of manually captioned video clips. Creating these datasets through human annotation is prohibitively expensive, time-consuming, subjective, and difficult to scale, which severely limits the scope and performance of existing models.

The article demonstrates that highly capable joint video-text representations can be trained entirely without manual annotations by using massive amounts of readily available instructional web videos paired with automatic speech transcriptions.

To achieve this, the authors collected HowTo100M, a dataset comprising 136 million video clips from 1.22 million narrated YouTube videos spanning over 23,000 physical tasks across 12 distinct categories. Using pre-extracted visual features and word embeddings, they trained a joint embedding model designed to map video clips and text into a shared semantic space. Crucially, the training process used an intra-video negative sampling strategy, pairing clips with incorrect narrations from the exact same video to force the model to focus on subtle action details rather than generic background cues.

The core findings show substantial performance and efficiency gains. First, the off-the-shelf model trained on HowTo100M establishes new state-of-the-art benchmarks on instructional video tasks: it achieved a 33.6% average recall for action step localization on the CrossTask benchmark (surpassing fully supervised baselines) and delivered top retrieval performance on cooking videos (YouCook2). Second, the model transfers remarkably well to non-instructional domains, such as generic web clips (MSR-VTT) and movie clips (LSMDC). Third, fine-tuning the pre-trained model on just 20% of the MSR-VTT dataset matched prior state-of-the-art performance trained on 100% of the data. Finally, empirical evaluations showed that model performance steadily improved as data volume grew, with no observed saturation at scale.

These results demonstrate that web-scale weakly supervised pre-training offers an effective, low-cost alternative to manual labeling, reducing the human annotation required to build strong domain-specific video models by up to 80%. Pre-training on large-scale instructional video consistently yields better downstream models across diverse video domains compared to training from scratch or pre-training on smaller datasets.

Organizations developing video search, retrieval, or activity analysis tools should adopt large-scale narrated video pre-training as a foundational base before fine-tuning on specialized downstream tasks. Rather than investing heavily in extensive manual captioning programs, teams can maximize efficiency by curating small, high-quality target datasets for domain-specific fine-tuning.

These conclusions come with certain caveats. Automatically transcribed narrations are inherently noisy; manual inspection showed that only about 51% of narrated clips strictly depict the spoken object or action. Additionally, because the source data focuses on physical how-to tasks, zero-shot performance on drastically different video formats, such as cinema, remains limited without fine-tuning. Despite this noise, the authors express high confidence that massive scale compensates for imperfect alignment, establishing a highly reliable and cost-effective methodology for video-language modeling.

arXiv: 1906.03327
Cover for HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Abstract

Learning text-video embeddings usually requires a dataset of video clips with manually provided captions. However, such datasets are expensive and time consuming to create and therefore difficult to obtain on a large scale. In this work, we propose instead to learn such embeddings from video data with readily available natural language annotations in the form of automatically transcribed narrations. The contributions of this work are three-fold. First, we introduce HowTo100M: a large-scale dataset of 136 million video clips sourced from 1.22M narrated instructional web videos depicting humans performing and describing over 23k different visual tasks. Our data collection procedure is fast, scalable and does not require any additional manual annotation. Second, we demonstrate that a text-video embedding trained on this data leads to state-of-the-art results for text-to-video retrieval and action localization on instructional video datasets such as YouCook2 or CrossTask. Finally, we show that this embedding transfers well to other domains: fine-tuning on generic Youtube videos (MSR-VTT dataset) and movies (LSMDC dataset) outperforms models trained on these datasets alone. Our dataset, code and models will be publicly available at: this http URL.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 The HowTo100M dataset
  • 3.1 Data collection
  • 3.2 Paired video clips and captions
  • 4 Text-video joint embedding model
  • 5 Experiments
  • 5.1 Implementation details
  • 5.2 Datasets and evaluation setups
  • 5.3 Study of negative pair sampling strategy
  • 5.4 Scale matters
  • 5.5 Comparison with state-of-the-art
  • 5.6 Cross-dataset fine-tuning evaluation
  • 5.7 Qualitative results
  • 6 Conclusion
  • References
  • A Additional details of the HowTo100M dataset
  • B Ranking loss implementation details
  • C Sampling strategy for positive pairs

Knowls

  1. Knowl 1 — HowTo100M Dataset Definition and Collection Protocol

    definition

    HowTo100M is a large-scale video dataset comprising 136.6 million video clips extracted from 1.221 million narrated instructional videos on YouTube, totaling approximately 134,472 hours (15 years) of video content across 23,611 distinct visual tasks.

    Category Tasks Videos Clips
    Food and Entertaining 11,504 497k 54.4M
    Home and Garden 5,068 270k 29.5M
    Hobbies and Crafts 4,273 251k 29.8M
    Cars Other Vehicles 810 68k 7.8M
    Pets and Animals 552 31k 3.5M
    Holidays and Traditions 411 27k 3.0M
    Personal Care and Style 181 16k 1.6M
    Sports and Fitness 205 16k 2.0M
    Health 172 15k 1.7M
    Education and Communications 239 15k 1.6M
    Arts and Entertainment 138 10k 1.2M
    Computers and Electronics 58 5k 0.6M
    Total 23,611 1.22M 136.6M

    The dataset is constructed via the following pipeline:

    1. Task Selection: 23,611 physical visual tasks are harvested from WikiHow across 12 high-level categories (excluding non-physical categories such as Relationships and Finance and Business). Tasks are filtered semi-automatically by restricting the root verb to physical interactions (e.g., make, build, change) while removing non-physical verbs (e.g., be, feel).
    2. Video Retrieval and Filtering: For each task, search queries formed as how to <task name> retrieve the top 200 YouTube search results. Videos are required to have English subtitles (user-uploaded, automatic speech recognition (ASR), or translated via YouTube API). Videos with fewer than 100 views, fewer than 100 spoken words, or durations exceeding 2,000 seconds are removed.
    3. Clip-Caption Pairing: Subtitle lines are paired directly with the corresponding video time interval. On average, each video yields 110 clip-caption pairs, with an average clip duration of 4 seconds and an average caption length of 4 non-stop words. Manual inspection indicates that in 51% of pairs, at least one mentioned object or action is visually visible in the corresponding clip.
  2. Knowl 2 — Gated Non-Linear Joint Text-Video Embedding Architecture

    model/method

    The text-video embedding model maps a pre-extracted video clip feature vector v∈Rdvv \in \mathbb{R}^{d_v} and a caption feature vector c∈Rdcc \in \mathbb{R}^{d_c} into a common dd-dimensional embedding space (dv=dc=d=4096d_v = d_c = d = 4096, comprising approximately 67 million trainable parameters) using non-linear gating functions.

    The mapping functions f:Rdv→Rdf: \mathbb{R}^{d_v} \to \mathbb{R}^d for video and g:Rdc→Rdg: \mathbb{R}^{d_c} \to \mathbb{R}^d for captions are defined as:

    f(v)=(W1vv+b1v)∘σ(W2v(W1vv+b1v)+b2v)f(v) = \left( W_1^v v + b_1^v \right) \circ \sigma\left( W_2^v \left( W_1^v v + b_1^v \right) + b_2^v \right)

    g(c)=(W1cc+b1c)∘σ(W2c(W1cc+b1c)+b2c)g(c) = \left( W_1^c c + b_1^c \right) \circ \sigma\left( W_2^c \left( W_1^c c + b_1^c \right) + b_2^c \right)

    where:

    • W1v∈Rd×dvW_1^v \in \mathbb{R}^{d \times d_v} and W1c∈Rd×dcW_1^c \in \mathbb{R}^{d \times d_c} are linear projection matrices,
    • W2v,W2c∈Rd×dW_2^v, W_2^c \in \mathbb{R}^{d \times d} are gating projection matrices,
    • b1v,b1c,b2v,b2c∈Rdb_1^v, b_1^c, b_2^v, b_2^c \in \mathbb{R}^d are bias vectors,
    • σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} is the element-wise sigmoid activation function acting as a context-gating mechanism with values in [0,1][0, 1],
    • ∘\circ denotes the element-wise Hadamard product.

    The compatibility between a video clip VV and a caption CC is computed as the cosine similarity between their normalized embeddings:

    s(V,C)=⟨f(v),g(c)⟩∥f(v)∥2∥g(c)∥2s(V, C) = \frac{\langle f(v), g(c) \rangle}{\|f(v)\|_2 \|g(c)\|_2}

  3. Knowl 3 — Max-Margin Ranking Loss with Intra-Video Negative Re-Weighting

    model/method

    The joint embedding is trained using a bi-directional max-margin ranking loss over mini-batches of paired clips and captions. To force the representation to distinguish between fine-grained visual actions occurring within the same video rather than generic background features, negative samples are sampled both from within the same video (intra-video negatives) and across different videos (inter-video negatives).

    A training mini-batch B\mathcal{B} is created by sampling v=32v = 32 unique YouTube videos and drawing k=64k = 64 clip-caption pairs with replacement from each, giving a total batch size of b=vk=2048b = vk = 2048 pairs. For every anchor pair i∈Bi \in \mathcal{B}, all remaining items in the batch form the negative set N(i)=B∖{i}\mathcal{N}(i) = \mathcal{B} \setminus \{i\}.

    To decouple the ratio of intra-video to inter-video negative loss contributions from the choice of vv and kk, a target intra-video sampling probability p∈[0,1]p \in [0, 1] (set to p=0.5p = 0.5) is enforced using sample-pair weights αi,j\alpha_{i,j}:

    αi,j={pk(v−1)(1−p)(k−1)if clip i and caption j originate from the same video,1otherwise.\alpha_{i,j} = \begin{cases} \dfrac{p k (v-1)}{(1-p)(k-1)} & \text{if clip } i \text{ and caption } j \text{ originate from the same video,} \\[10pt] 1 & \text{otherwise.} \end{cases}

    The total mini-batch loss minimized via the Adam optimizer (learning rate 10−410^{-4}, margin δ=0.1\delta = 0.1) is:

    L=∑i∈B∑j∈N(i)αi,j[max⁡(0,δ+s(Vi,Cj)−s(Vi,Ci))+max⁡(0,δ+s(Vj,Ci)−s(Vi,Ci))]\mathcal{L} = \sum_{i \in \mathcal{B}} \sum_{j \in \mathcal{N}(i)} \alpha_{i,j} \left[ \max(0, \delta + s(V_i, C_j) - s(V_i, C_i)) + \max(0, \delta + s(V_j, C_i) - s(V_i, C_i)) \right]

  4. Knowl 4 — Video and Text Feature Extraction Pipeline

    experimental setup

    Input representations for video clips and spoken narrations are generated as follows:

    1. Visual Features (v∈R4096v \in \mathbb{R}^{4096}):

      • 2D visual cues are computed by extracting 2048-dimensional features at 1 frame per second using a ResNet-152 model pre-trained on ImageNet.
      • 3D motion and temporal cues are computed by extracting 2048-dimensional features at 1.5 representations per second (using 16-frame windows) using a ResNeXt-101 3D-CNN pre-trained on the Kinetics-400 dataset.
      • Features across the duration of each clip are aggregated via temporal max-pooling and concatenated, producing a single 4096-dimensional visual vector vv.
    2. Text Features (c∈R4096c \in \mathbb{R}^{4096}):

      • Transcribed subtitle text is filtered to remove standard English stop-words.
      • Word tokens are embedded using pre-trained GoogleNews Word2Vec embeddings (300 dimensions).
      • A shallow 1D Convolutional Neural Network (1D-CNN) processes the sequence of word embeddings to output a 4096-dimensional caption feature vector cc.
    3. Training Time: Extraction of static features enables full training of the joint embedding model on HowTo100M in under 3 days on a single Nvidia Tesla P100 GPU.

  5. Knowl 5 — Action Step Localization Performance on CrossTask

    empirical result

    The text-video embedding trained solely on HowTo100M without fine-tuning establishes state-of-the-art action step localization on the CrossTask benchmark (18 tasks, 2.7k videos). Frame-level step assignments are obtained by computing cosine similarity between each video frame embedding and natural language action step descriptions. Performance is measured by average recall (percentage of assigned steps falling into ground-truth intervals).

    Method Average Recall (%)
    Alayrac et al. (Weakly-supervised) 13.3
    Zhukov et al. (Weakly-supervised) 22.4
    Fully-supervised Upper-Bound 31.6
    Ours (trained on HowTo100M only) 33.6

    The zero-shot HowTo100M embedding surpasses the previous weakly supervised state of the art by +11.2% absolute recall and exceeds the fully supervised upper-bound model (31.6%) trained on ground-truth segmented action annotations.

  6. Knowl 6 — Text-to-Video Retrieval Performance on YouCook2

    empirical result

    Evaluation of text-to-video clip retrieval on the 3.5k validation clips of the YouCook2 instructional cooking dataset compares the HowTo100M pre-trained embedding (with and without fine-tuning) against baseline methods. Evaluation metrics are Recall@k (R@1, R@5, R@10, higher is better) and Median Rank (Median R, lower is better).

    Method Training Set R@1 (%) R@5 (%) R@10 (%) Median R
    Random None 0.03 0.15 0.3 1675
    HGLMM FV CCA YouCook2 4.6 14.3 21.6 75
    Ours YouCook2 4.2 13.7 21.5 65
    Ours HowTo100M 6.1 17.3 24.8 46
    Ours (PT: HowTo100M, FT: YouCook2) YouCook2 8.2 24.5 35.3 24

    Direct zero-shot application of the model trained on HowTo100M outperforms the same architecture trained directly on YouCook2 (R@10 of 24.8% vs. 21.5%). Fine-tuning the HowTo100M model on YouCook2 achieves 35.3% R@10, a +13.7% gain over previous methods.

  7. Knowl 7 — Text-to-Video Retrieval Transfer to Generic YouTube Videos (MSR-VTT)

    empirical result

    Evaluation of text-to-video clip retrieval on the 1,000-clip test set of MSR-VTT demonstrates the cross-domain transfer capability of HowTo100M pre-training onto generic web videos.

    Method Training Set R@1 (%) R@5 (%) R@10 (%) Median R
    Random None 0.1 0.5 1.0 500
    C+LSTM+SA+FC7 MSR-VTT 4.2 12.9 19.9 55
    VSE-LSTM MSR-VTT 3.8 12.7 17.1 66
    SNUVL MSR-VTT 3.5 15.9 23.8 44
    Kaufman et al. MSR-VTT 4.7 16.6 24.1 41
    CT-SAN MSR-VTT 4.4 16.6 22.3 35
    JSFusion MSR-VTT 10.2 31.2 43.2 13
    Ours HowTo100M 7.5 21.2 29.6 38
    Ours MSR-VTT 12.1 35.0 48.0 12
    Ours (PT: HowTo100M, FT: MSR-VTT) MSR-VTT 14.9 40.2 52.8 9

    Zero-shot HowTo100M embeddings outperform several supervised models trained directly on MSR-VTT. Pre-training on HowTo100M followed by fine-tuning on MSR-VTT outperforms the supervised state-of-the-art JSFusion (52.8% vs. 43.2% R@10). Fine-tuning the pre-trained model on only 20% of MSR-VTT labeled training data matches the accuracy of JSFusion trained on 100% of MSR-VTT data.

  8. Knowl 8 — Text-to-Video Retrieval Transfer to Movie Clips (LSMDC)

    empirical result

    Evaluation of text-to-video clip retrieval on the 1,000-clip test set of the Large Scale Movie Description Challenge (LSMDC) assesses transfer across the substantial domain gap between instructional YouTube videos and movies with script/audio descriptions.

    Method Training Set R@1 (%) R@5 (%) R@10 (%) Median R
    Random None 0.1 0.5 1.0 500
    C+LSTM+SA+FC7 LSMDC 4.3 12.6 18.9 98
    VSE-LSTM LSMDC 3.1 10.4 16.5 79
    SNUVL LSMDC 3.6 14.7 23.9 50
    Kaufman et al. LSMDC 4.7 15.9 23.4 64
    CT-SAN LSMDC 4.5 14.1 20.9 67
    JSFusion LSMDC 9.1 21.2 34.1 36
    Ours HowTo100M 4.0 9.8 14.0 137
    Ours LSMDC 7.2 18.3 25.0 44
    Ours (PT: HowTo100M, FT: LSMDC) LSMDC 7.1 19.6 27.9 40

    Fine-tuning a HowTo100M pre-trained embedding on LSMDC achieves 27.9% R@10, outperforming direct training on LSMDC alone (25.0% R@10).

  9. Knowl 9 — Impact of Data Scale on Downstream Performance

    empirical result

    Evaluating the text-video embedding trained on subsets of HowTo100M demonstrates that downstream performance scales monotonically with the number of training videos without signs of saturation.

    Subsets were constructed by varying the allowable YouTube search rank per task: top 2 (15k videos), top 3 (28k videos), top 5 (52k videos), top 10 (104k videos), top 20 (197k videos), top 40 (364k videos), top 80 (648k videos), and top 200 (all 1.22M videos). Across all four downstream benchmarks (CrossTask average recall, LSMDC R@10, MSR-VTT R@10, and YouCook2 R@10), metrics improve consistently as data scale increases from 15k to 1.22M videos.

  10. Knowl 10 — Ablations on Intra-Video Negative Sampling and Positive Pair Filtering

    empirical result

    Ablation experiments evaluate the impact of negative sampling strategies and positive pair filtering techniques on retrieval (R@10 on MSR-VTT, LSMDC, YouCook2) and step localization (average recall on CrossTask):

    1. Intra-Video Negative Sampling: Mining negative clips and captions from within the same video is essential for fine-grained action discrimination:
    Negative Sampling Strategy MSR-VTT (R@10) LSMDC (R@10) YouCook2 (R@10) CrossTask (Avg Recall %)
    No intra-negatives 30.1 12.3 18.1 25.7
    With intra-negatives 29.6 14.0 24.8 33.6
    1. Positive Pair Filtering (Max-Pooling): Attempting to filter out noisy positive pairs by retaining only the proportion r∈[0.2,1.0]r \in [0.2, 1.0] of highest-scoring positive pairs per video under current model parameters does not improve performance; the unconstrained model (r=1.0r = 1.0) performs best (MSR-VTT: 29.6%, LSMDC: 14.0%, YouCook2: 24.8% vs. r=0.5r = 0.5: 25.2%, 12.6%, 23.5%). The model demonstrates natural robustness to weak alignment noise when trained over massive video data.

Coverage note — None was omitted; all contributed datasets, model formulations, loss weightings, feature pipelines, and primary experimental benchmarks (CrossTask, YouCook2, MSR-VTT, LSMDC, scaling, and ablations) are fully covered.

References

  1. 1.Project webpage. https://www.di.ens.fr/willow/research/howto100m/, 2019. 1, 8, 9
  2. 2.J.-B. Alayrac, P. Bojanowski, N. Agrawal, I. Laptev, J. Sivic, and S. Lacoste-Julien. Unsupervised learning from narrated instruction videos. In CVPR, 2016. 2, 6, 7
  3. 3.J.-B. Alayrac, J. Sivic, I. Laptev, and S. Lacoste-Julien. Joint discovery of object states and manipulation actions. In ICCV, 2017. 2
  4. 4.J. Carreira and A. Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR, 2017. 5
  5. 5.K. Chen, H. Song, C. Change Loy, and D. Lin. Discover and learn new objects from documentaries. In CVPR, 2017. 2
  6. 6.M. Chowdhury, P. Rameswar, E. Papalexakis, and A. Roy-Chowdhury. Webly supervised joint embedding for cross-modal image-text retrieval. In ACM International Conference on Multimedia, 2018. 2
  7. 7.D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al. Scaling egocentric vision: The epic-kitchens dataset. In ECCV, 2018. 2
  8. 8.J. Dong, X. Li, C. Xu, S. Ji, Y. He, G. Yang, and X. Wang. Dual encoding for zero-example video retrieval. In CVPR, 2019. 2
  9. 9.A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach. Multimodal compact bilinear pooling for visual question answering and visual grounding. In EMNLP, pages 457–468, 2016. 2
  10. 10.Y. Gong, Q. Ke, M. Isard, and S. Lazebnik. A multi-view embedding space for modeling internet images, tags, and their semantics. IJCV, 2014. 2
  11. 11.Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik. Improving image-sentence embeddings using large weakly annotated photo collections. In ECCV, 2014. 2
  12. 12.K. Hara, H. Kataoka, and Y. Satoh. Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet? In CVPR, 2018. 5
  13. 13.D. Harwath, A. Recasens, D. Sur'ıs, G. Chuang, A. Torralba, and J. Glass. Jointly discovering visual objects and spoken words from raw sensory input. In ECCV, 2018. 2
  14. 14.K. He, X. Zhang, S. Ren, and J. Sun. Deep Residual Learning for Image Recognition. In CVPR, 2016. 5
  15. 15.L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell. Localizing moments in video with natural language. ICCV, 2017. 1, 2, 5, 12
  16. 16.D.-A. Huang, L. Fei-Fei, and J. C. Niebles. Connectionist temporal modeling for weakly supervised action labeling. In ECCV, 2016. 2
  17. 17.D.-A. Huang, J. J. Lim, L. Fei-Fei, and J. C. Niebles. Unsupervised visual-linguistic reference resolution in instructional videos. In CVPR, 2017. 2
  18. 18.D.-A. Huang, V. Ramanathan, D. Mahajan, L. Torresani, M. Paluri, L. Fei-Fei, and J. C. Niebles. Finding "it": Weakly-supervised reference-aware visual grounding in instructional video. In CVPR, 2018. 2
  19. 19.K. L. Jacob Devlin, Ming-Wei Chang and K. Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. preprint, 2018. 3
  20. 20.J. Johnson, A. Karpathy, and L. Fei-Fei. Densecap: Fully convolutional localization networks for dense captioning. In CVPR, 2016. 2
  21. 21.A. Karpathy, A. Joulin, and F. F. F. Li. Deep fragment embeddings for bidirectional image sentence mapping. In NIPS, 2014. 5
  22. 22.D. Kauman, G. Levi, T. Hassner, and L. Wolf. Temporal tessellation: A unified approach for video analysis. In ICCV, 2017. 7, 8
  23. 23.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 5
  24. 24.R. Kiros, R. Salakhutdinov, and R. S. Zemel. Unifying visual-semantic embeddings with multimodal neural language models. TACL, 2014. 7, 8
  25. 25.B. Klein, G. Lev, G. Sadeh, and L. Wolf. Associating neural word embeddings with deep image representations using fisher vectors. In CVPR, 2015. 1, 2, 7
  26. 26.R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles. Dense-captioning events in videos. In ICCV, 2017. 2
  27. 27.Y. Li, Y. Song, L. Cao, J. Tetreault, L. Goldberg, A. Jaimes, and J. Luo. TGIF: A New Dataset and Benchmark on Animated GIF Description. In CVPR, 2016. 2
  28. 28.D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. van der Maaten. Exploring the limits of weakly supervised pretraining. In ECCV, 2018. 3
  29. 29.M. Malinowski, M. Rohrbach, and M. Fritz. Ask your neurons: A neural-based approach to answering questions about images. In ICCV, 2015. 2
  30. 30.J. Malmaud, J. Huang, V. Rathod, N. Johnston, A. Rabinovich, and K. Murphy. What's cookin'? interpreting cooking videos using text, speech and vision. NAACL, 2015. 2
  31. 31.A. Miech, I. Laptev, and J. Sivic. Learnable pooling with context gating for video classification. arXiv preprint arXiv:1706.06905, 2017. 5
  32. 32.A. Miech, I. Laptev, and J. Sivic. Learning a Text-Video Embedding from Incomplete and Heterogeneous Data. arXiv:1804.02516, 2018. 1, 2, 4, 5
  33. 33.A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic. Howto100M: Learning a text-video embedding by watching hundred million narrated video clips. arXiv preprint arXiv:1906.03327, 2019. 4, 5
  34. 34.T. Mikolov, K. Chen, G. Corrado, and J. Dean. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781, 2013. 5
  35. 35.N. C. Mithun, J. Li, F. Metze, and A. K. Roy-Chowdhury. Learning joint embedding with multimodal cues for cross-modal video-text retrieval. In Proceedings of the 2018 ACM on International Conference on Multimedia Retrieval. ACM, 2018. 2
  36. 36.P. Pan, Z. Xu, Y. Yang, F. Wu, and Y. Zhuang. Hierarchical recurrent neural encoder for video representation with application to captioning. In CVPR, pages 1029–1038, 2016. 1, 2
  37. 37.Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui. Jointly modeling embedding and translation to bridge video and language. In CVPR, 2016. 1, 2
  38. 38.B. A. Plummer, M. Brown, and S. Lazebnik. Enhancing video summarization via vision-language embedding. In CVPR, 2017. 1, 2
  39. 39.A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever. Improving Language Understandingby Generative Pre-Training. preprint, 2018. 3
  40. 40.A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. preprint, 2019. 3
  41. 41.A. Richard, H. Kuehne, and J. Gall. Weakly supervised action learning with rnn based fine-to-coarse modeling. In CVPR, 2017. 2
  42. 42.A. Richard, H. Kuehne, and J. Gall. Action sets: Weakly supervised action segmentation without ordering constraints. In CVPR, 2018. 2
  43. 43.A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele. A dataset for movie description. In CVPR, 2015. 2
  44. 44.A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele. Movie description. IJCV, 2017. 2, 5, 6
  45. 45.R. Sanabria, O. Caglayan, S. Palaskar, D. Elliott, L. Barrault, L. Specia, and F. Metze. How2: a large-scale dataset for multimodal language understanding. In Proceedings of the Workshop on Visually Grounded Interaction and Language (ViGIL). NeurIPS, 2018. 2
  46. 46.F. Sener and A. Yao. Unsupervised learning and segmentation of complex activities from video. In CVPR, 2018. 2
  47. 47.O. Sener, A. R. Zamir, S. Savarese, and A. Saxena. Unsupervised semantic parsing of video collections. In The IEEE International Conference on Computer Vision (ICCV), December 2015. 2
  48. 48.G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta. Hollywood in homes: Crowdsourcing data collection for activity understanding. In European Conference on Computer Vision, 2016. 2
  49. 49.C. Sun, A. Shrivastava, S. Singh, and A. Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In ICCV, 2017. 3
  50. 50.Y. Tang, D. Ding, Y. Rao, Y. Zheng, D. Zhang, L. Zhao, J. Lu, and J. Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In CVPR, 2019. 2
  51. 51.M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler. Movieqa: Understanding stories in movies through question-answering. In CVPR, 2016. 1, 2
  52. 52.A. Torabi, C. Pal, H. Larochelle, and A. Courville. Using descriptive video services to create a large data source for video annotation research. arXiv preprint arXiv:1503.01070, 2015. 2
  53. 53.A. Torabi, N. Tandon, and L. Sigal. Learning language-visual embedding for movie understanding with natural-language. arXiv preprint arXiv:1609.08124, 2016. 7, 8
  54. 54.L. Wang, Y. Li, J. Huang, and S. Lazebnik. Learning two-branch neural networks for image-text matching tasks. PAMI, 2018. 1, 2, 5
  55. 55.L. Wang, Y. Li, and S. Lazebnik. Learning deep structure-preserving image-text embeddings. In CVPR, pages 5005–5013, 2016. 1, 2, 5
  56. 56.X. Wang, J. Wu, D. Zhang, Y. Su, and W. Y. Wang. Learning to compose topic-aware mixture of experts for zero-shot video captioning. In AAAI, 2018. 2
  57. 57.C.-Y. Wu, R. Manmatha, A. J. Smola, and P. Kr¨ahenb¨uhl. Sampling matters in deep embedding learning. ICCV, 2017. 2
  58. 58.J. Xu, T. Mei, T. Yao, and Y. Rui. Msr-vtt: A large video description dataset for bridging video and language. In CVPR, 2016. 1, 2, 5, 6
  59. 59.R. Xu, C. Xiong, W. Chen, and J. J. Corso. Jointly modeling deep video and compositional text to bridge vision and language in a unified framework. In AAAI, volume 5, page 6, 2015. 1, 2
  60. 60.Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo. Image captioning with semantic attention. In CVPR, pages 4651–4659, 2016. 2
  61. 61.H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu. Video paragraph captioning using hierarchical recurrent neural networks. In CVPR, pages 4584–4593, 2016. 1, 2
  62. 62.S.-I. Yu, L. Jiang, and A. Hauptmann. Instructional videos for unsupervised harvesting and learning of action examples. In ACM, 2014. 2
  63. 63.Y. Yu, J. Kim, and G. Kim. A joint sequence fusion model for video question answering and retrieval. In ECCV, 2018. 1, 2, 6, 7, 8
  64. 64.Y. Yu, H. Ko, J. Choi, and G. Kim. Video captioning and retrieval models with semantic attention. In ECCV LSMDC2016 Workshop, 2016. 5, 7, 8
  65. 65.Y. Yu, H. Ko, J. Choi, and G. Kim. End-to-end concept word detection for video captioning, retrieval, and question answering. In CVPR, 2017. 7, 8
  66. 66.L. Zhou, X. Chenliang, and J. J. Corso. Towards automatic learning of procedures from web instructional videos. In AAAI, 2018. 2
  67. 67.L. Zhou, C. Xu, and J. J. Corso. Towards automatic learning of procedures from web instructional videos. In AAAI, 2018. 2, 5, 6, 7
  68. 68.D. Zhukov, J.-B. Alayrac, R. G. Cinbis, D. Fouhey, I. Laptev, and J. Sivic. Cross-task weakly supervised learning from instructional videos. In CVPR, 2019. 2, 5, 6, 7

Citation

MLA
Miech, A., et al. “HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips”. arXiv, 2019, http://arxiv.org/abs/1906.03327v2.
APA
Miech, A., Zhukov, D., Alayrac, J.-B., Tapaswi, M., Laptev, I., & Sivic, J. (2019). HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips. arXiv. http://arxiv.org/abs/1906.03327v2
Chicago
Miech, A., D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic. 2019. “HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips”. arXiv. http://arxiv.org/abs/1906.03327v2.
Harvard
Miech, A. et al. (2019) “HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1906.03327v2.
Vancouver
1. Miech A, Zhukov D, Alayrac J-B, Tapaswi M, Laptev I, Sivic J (2019) HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips. arXiv

BibTeX

@article{miech2019howto100m,
  title = {HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips},
  author = {Miech, Antoine and Zhukov, Dimitri and Alayrac, Jean-Baptiste and Tapaswi, Makarand and Laptev, Ivan and Sivic, Josef},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1906.03327v2},
  eprint = {1906.03327}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE