HierVL: Learning Hierarchical Video-Language Embeddings

Kumar AshutoshRohit GirdharLorenzo TorresaniKristen Grauman

article2023CVPR87 citationsDistinguished Paper Award

Proposes a hierarchical video-language framework that jointly aligns short-term action clips with step-by-step descriptions and aggregated video features with abstract summaries to capture both immediate actions and long-term actor intent.

Listen

Artificial intelligence models have achieved human-level performance in static image analysis, but interpreting human activity in video remains a significant hurdle. Standard video-language systems connect short, seconds-long clips directly to immediate descriptions (such as "opening a tap"), failing to capture the overarching context, sequential dependencies, and broader human intent (such as "preparing dinner"). This gap limits the effectiveness of automated video analysis in high-value domains like robotics, augmented reality, and large-scale media retrieval.

The article demonstrates a novel hierarchical video-language learning framework, named HierVL. The primary objective is to evaluate whether jointly training models on both granular, immediate actions and high-level abstract summaries creates visual representations that simultaneously capture short-term movements and long-term goals without incurring prohibitive computational costs.

To achieve this, the authors developed a dual-layer training framework using the Ego4D dataset, which contains 3,670 hours of daily-life wearable camera video annotated with 3.85 million step-by-step narrations and 120,000 video-level summaries. The approach matches individual video clips to step-by-step descriptions at the child level, while aggregating clip features across entire videos using self-attention mechanisms to match high-level summary texts at the parent level. This joint contrastive process circumvents memory bottlenecks associated with processing long raw videos and prevents catastrophic forgetting across different levels of abstraction.

The evaluation produced several decisive findings. First, HierVL-SA (using self-attention aggregation) achieved 95.4% accuracy on long-term summary matching and 26.8% on temporal sequence ordering tests, beating prior state-of-the-art baselines like EgoVLP by more than 6% and achieving a 34% relative improvement in temporal ordering where baseline models scored at chance levels (20%). Second, the framework established new state-of-the-art results across diverse external benchmarks, including the Ego4D Long-Term Anticipation challenge (predicting the next 20 actions), Charades-Ego action recognition (reaching 33.8% mean average precision when fine-tuned), and EPIC-KITCHENS-100 multi-instance retrieval. Third, linear probe tests on HowTo100M video classification demonstrated strong transferability, reaching 64.6% accuracy compared to 53.4% for existing methods, while resisting the transfer overfitting seen in standard models.

These findings indicate that incorporating abstract human intent fundamentally enriches foundational visual representations. For organizations developing video analytics, robotics, or interactive AI assistants, adopting hierarchical modeling directly improves long-horizon task planning and action forecasting while maintaining low-level precision. Importantly, downstream implementations do not require text summaries during deployment, making the learned representations immediately usable for standard video tasks without operational overhead.

Organizations advancing video-understanding pipelines should adopt hierarchical pretraining strategies and incorporate high-level summary metadata where available. Technical teams should prioritize self-attention aggregation over basic average pooling when temporal ordering is critical to the application. Future work should investigate scaling this approach to larger non-egocentric video corpora and exploring more granular intermediate hierarchy levels.

Confidence in these findings is high given rigorous benchmarking across multiple established datasets and explicit ablation studies confirming the necessity of both summary supervision and hierarchical structure. However, readers should note that training relies on datasets with dual-level text annotations, and variations in computing hardware configurations can slightly affect absolute baseline reproductions.

arXiv: 2301.02311
Cover for HierVL: Learning Hierarchical Video-Language Embeddings

Abstract

Video-language embeddings are a promising avenue for injecting semantics into visual representations, but existing methods capture only short-term associations between seconds-long video clips and their accompanying text. We propose HierVL, a novel hierarchical video-language embedding that simultaneously accounts for both long-term and short-term associations. As training data, we take videos accompanied by timestamped text descriptions of human actions, together with a high-level text summary of the activity throughout the long video (as are available in Ego4D). We introduce a hierarchical contrastive training objective that encourages text-visual alignment at both the clip level and video level. While the clip-level constraints use the step-by-step descriptions to capture what is happening in that instant, the video-level constraints use the summary text to capture why it is happening, i.e., the broader context for the activity and the intent of the actor. Our hierarchical scheme yields a clip representation that outperforms its single-level counterpart as well as a long-term video representation that achieves SotA results on tasks requiring long-term video modeling. HierVL successfully transfers to multiple challenging downstream tasks (in EPIC-KITCHENS-100, Charades-Ego, HowTo100M) in both zero-shot and fine-tuned settings.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Technical Approach
  • 3.1. Hierarchical video annotations
  • 3.2. Hierarchical joint video and text embedding
  • 3.3. Efficient long-term features via aggregation
  • 3.4. Contrastive pretraining objective
  • 3.5. Training strategy
  • 3.6. Implementation Details
  • 4. Experiments
  • 4.1. Pretraining Evaluation
  • 4.2. Downstream Evaluation
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — HierVL Hierarchical Video-Language Embedding Framework

    model/method

    HierVL is a multi-modal representation learning framework that models human activity simultaneously at fine-grained (clip) and high-level (video) temporal granularities by projecting visual and textual representations into a shared embedding space.

    A hierarchically annotated video dataset is defined as DL={(Vi,Ni,Si)}i=1∣DL∣\mathcal{D}_L = \{(V_i, N_i, S_i)\}_{i=1}^{|\mathcal{D}_L|}, where:

    • Vi={vij}j=1∣Vi∣V_i = \{v_{ij}\}_{j=1}^{|V_i|} is a long video represented as an ordered sequence of short visual clips vijv_{ij} spanning a few seconds.
    • Ni={nij}j=1∣Ni∣N_i = \{n_{ij}\}_{j=1}^{|N_i|} is an ordered sequence of short-term textual narrations nijn_{ij}, each describing the atomic action in clip vijv_{ij}.
    • SiS_i is a global high-level abstractive text summary (typically 1–3 sentences) describing the overall activity, environment, and actor intent across the entire duration of video ViV_i.

    The model maps inputs into a shared embedding space using:

    • A visual clip encoder fv(v)f_v(v) producing short-term visual representations.
    • A text encoder fn(n)f_n(n) producing short-term text representations, and also encoding the summary text fn(S)f_n(S).
    • Long-term visual representations fV(Vi)=Agg({fv(vij)}j=1∣Vi∣)f_V(V_i) = \text{Agg}(\{f_v(v_{ij})\}_{j=1}^{|V_i|}) and long-term textual representations fN(Ni)=Agg({fn(nij)}j=1∣Ni∣)f_N(N_i) = \text{Agg}(\{f_n(n_{ij})\}_{j=1}^{|N_i|}), where Agg\text{Agg} is an aggregation function.

    For a similarity metric sim(⋅,⋅)\text{sim}(\cdot, \cdot), the embedding space satisfies two levels of matching constraints:

    1. Child-level matching constraint: For all (i1,j1)≠(i2,j2)(i_1, j_1) \neq (i_2, j_2), sim(fv(vi1j1),fn(ni1j1))>sim(fv(vi1j1),fn(ni1j2))\text{sim}\left(f_v(v_{i_1 j_1}), f_n(n_{i_1 j_1})\right) > \text{sim}\left(f_v(v_{i_1 j_1}), f_n(n_{i_1 j_2})\right)
    2. Parent-level matching constraints: For all video indices i≠ji \neq j, sim(fV(Vi),fn(Si))>sim(fV(Vi),fn(Sj))\text{sim}\left(f_V(V_i), f_n(S_i)\right) > \text{sim}\left(f_V(V_i), f_n(S_j)\right) sim(fN(Ni),fn(Si))>sim(fN(Ni),fn(Sj))\text{sim}\left(f_N(N_i), f_n(S_i)\right) > \text{sim}\left(f_N(N_i), f_n(S_j)\right)
  2. Knowl 2 — HierVL Feature Aggregators: HierVL-SA and HierVL-Avg

    model/method

    To circumvent the computational cost and memory overflows of directly encoding entire long videos, HierVL computes long-term visual and textual representations by aggregating short-term features. Given KK short-term feature vectors (in practice K=16K = 16 uniformly sampled from the entire video or narration sequence), long-term visual and textual representations are formed as: fV(Vi)=Agg({fv(vij)}j=1K),fN(Ni)=Agg({fn(nij)}j=1K)f_V(V_i) = \text{Agg}\left(\{f_v(v_{ij})\}_{j=1}^{K}\right), \quad f_N(N_i) = \text{Agg}\left(\{f_n(n_{ij})\}_{j=1}^{K}\right)

    Two aggregation mechanisms Agg(⋅)\text{Agg}(\cdot) are used across both visual and text modalities:

    1. HierVL-SA (Self-Attention Aggregator): Implemented as a 6-layer self-attention transformer block with learnable positional encodings. The positional encodings allow the aggregator to capture temporal sequence order and cross-clip temporal dependencies across the full video.
    2. HierVL-Avg (Average Pooling Aggregator): A parameter-free aggregator that computes the mean of the KK short-term representation vectors. This aggregator assigns equal weight to all short-term clips and is permutation invariant (does not preserve temporal ordering).
  3. Knowl 3 — HierVL Hierarchical Contrastive Pretraining Objective

    equation

    HierVL pretraining optimizes a joint multi-level contrastive loss consisting of a child-level loss Lchild\mathcal{L}_{\text{child}} and a parent-level loss Lparent\mathcal{L}_{\text{parent}}.

    The child-level loss aligns short-term visual clips viv_i with step-by-step narrations nin_i: Lchild=−1∣B~∣∑i∈B~log⁡(∑j∈P~iexp⁡(fv(vi)Tfn(nj))∑j∈B~exp⁡(fv(vi)Tfn(nj)))\mathcal{L}_{\text{child}} = -\frac{1}{|\tilde{\mathcal{B}}|} \sum_{i \in \tilde{\mathcal{B}}} \log \left( \frac{\sum_{j \in \tilde{\mathcal{P}}_i} \exp\left(f_v(v_i)^T f_n(n_j)\right)}{\sum_{j \in \tilde{\mathcal{B}}} \exp\left(f_v(v_i)^T f_n(n_j)\right)} \right) where B~\tilde{\mathcal{B}} denotes the batch of short-term visual-text pairs, and P~i\tilde{\mathcal{P}}_i is the set of action-aware positive samples (distinct visual instances representing identical actions). Unlike standard EgoNCE, HierVL omits penalties on temporally close distinct actions to avoid penalizing atomic actions sharing the same overarching wearer intent.

    The parent-level loss aligns long-term aggregated representations with the abstractive summary text SiS_i: Lparent=LparentSV+LparentSN\mathcal{L}_{\text{parent}} = \mathcal{L}_{\text{parent}}^{SV} + \mathcal{L}_{\text{parent}}^{SN} where: LparentSV=−1∣B~∣∑i∈B~log⁡(∑j∈P~iexp⁡(fV(Vi)Tfn(Sj))∑j∈B~exp⁡(fV(Vi)Tfn(Sj)))\mathcal{L}_{\text{parent}}^{SV} = -\frac{1}{|\tilde{\mathcal{B}}|} \sum_{i \in \tilde{\mathcal{B}}} \log \left( \frac{\sum_{j \in \tilde{\mathcal{P}}_i} \exp\left(f_V(V_i)^T f_n(S_j)\right)}{\sum_{j \in \tilde{\mathcal{B}}} \exp\left(f_V(V_i)^T f_n(S_j)\right)} \right) LparentSN=−1∣B~∣∑i∈B~log⁡(∑j∈P~iexp⁡(fN(Ni)Tfn(Sj))∑j∈B~exp⁡(fN(Ni)Tfn(Sj)))\mathcal{L}_{\text{parent}}^{SN} = -\frac{1}{|\tilde{\mathcal{B}}|} \sum_{i \in \tilde{\mathcal{B}}} \log \left( \frac{\sum_{j \in \tilde{\mathcal{P}}_i} \exp\left(f_N(N_i)^T f_n(S_j)\right)}{\sum_{j \in \tilde{\mathcal{B}}} \exp\left(f_N(N_i)^T f_n(S_j)\right)} \right) For a summary SiS_i, negative visual and textual samples are chosen from outside the temporal span of SiS_i.

  4. Knowl 4 — Joint Hierarchical Alternating Training Strategy

    algorithm

    HierVL jointly optimizes the visual encoder fvf_v, the text encoder fnf_n, and the long-term aggregator Agg\text{Agg} using an alternating batch training scheme. This joint procedure prevents catastrophic forgetting that occurs when training sequentially on clip-level and video-level data, while ensuring that the short-term encoders fvf_v and fnf_n are directly shaped by long-term intent.

    Input: Training dataset DL={(Vi,Ni,Si)}\mathcal{D}_L = \{(V_i, N_i, S_i)\}, interval parameter m=5m = 5, batch sizes, learning rate η\eta
    Output: Trained encoders fv,fnf_v, f_n and aggregator Agg\text{Agg}
    Initialize parameters of fvf_v, fnf_n, and Agg\text{Agg}
    for each pretraining iteration do
        for step = 1 to mm do
            Sample a batch of short-term visual-narration pairs (v,n)(v, n) from different videos
            Compute short-term representations fv(v)f_v(v) and fn(n)f_n(n)
            Compute child loss Lchild\mathcal{L}_{\text{child}}
            Update parameters of fvf_v and fnf_n via gradient descent on Lchild\mathcal{L}_{\text{child}}
        end for
        Sample a batch of long-term instances, where each instance consists of K=16K=16 clips from the same video ViV_i with summary SiS_i
        Compute short-term clip features {fv(vij)}j=1K\{f_v(v_{ij})\}_{j=1}^K and narration features {fn(nij)}j=1K\{f_n(n_{ij})\}_{j=1}^K
        Compute aggregated features fV(Vi)=Agg({fv(vij)})f_V(V_i) = \text{Agg}(\{f_v(v_{ij})\}) and fN(Ni)=Agg({fn(nij)})f_N(N_i) = \text{Agg}(\{f_n(n_{ij})\})
        Compute summary text feature fn(Si)f_n(S_i)
        Compute parent loss Lparent=LparentSV+LparentSN\mathcal{L}_{\text{parent}} = \mathcal{L}_{\text{parent}}^{SV} + \mathcal{L}_{\text{parent}}^{SN}
        Update parameters of Agg\text{Agg}, fvf_v, and fnf_n via gradient descent on Lparent\mathcal{L}_{\text{parent}}
    end for
    return fv,fn,Aggf_v, f_n, \text{Agg}
  5. Knowl 5 — Pretraining Evaluation on EgoMCQ, SummaryMCQ, and ShuffleMCQ

    data/table

    Pretraining performance was evaluated on Ego4D using three multiple-choice tasks (chance accuracy is 20.0% across all tasks):

    • EgoMCQ: Clip-level text-to-video matching against 5 candidate clips from different videos (Inter-video) or the same video (Intra-video).
    • SummaryMCQ: Video-level text-to-video matching between an overarching summary and 5 candidate multi-minute videos spanning the summary duration.
    • ShuffleMCQ: Temporal order reasoning task where the model matches a summary to the correct video among 5 candidates, where 4 candidates have their constituent clips randomly shuffled in temporal order.
    Method Joint train Hier Summ Aggregation Summ MCQ Shuffle MCQ EgoMCQ
    Inter-video Intra-video
    EgoVLP — — — — — — 90.6 57.2
    EgoVLP (reproduced) — X X X 89.0 20.0 90.1 54.0
    HierVL-Avg (Ours) ✓ ✓ ✓ Average 95.2 20.0 90.3 53.1
    HierVL-SA (Ours) ✓ ✓ ✓ Self-attention 95.4 26.8 90.5 52.4
    HierVL-w/o Joint X X ✓ X 89.8 24.2 72.0 29.4
    HierVL-w/o Hier ✓ X ✓ X 93.7 20.0 90.7 50.5
    HierVL-w/o Summ ✓ ✓ X Self-attention 20.0 22.1 90.8 52.1
    HierVL-w/o Summ ↔\leftrightarrow Narr ✓ ✓ ✓ Self-attention 94.7 26.1 90.4 50.0

    HierVL-SA achieves a 6.4% gain on SummaryMCQ over EgoVLP and 26.8% accuracy on ShuffleMCQ (a 34% relative improvement over chance/EgoVLP/HierVL-Avg at 20.0%), confirming that self-attention aggregation captures temporal ordering dependencies. The ablations demonstrate that: (1) training without joint short/long-term optimization (HierVL-w/o Joint) causes catastrophic forgetting on clip-level EgoMCQ; (2) associating summaries to single clips without hierarchical aggregation (HierVL-w/o Hier) degrades video-level performance; (3) aggregating narrations without summary supervision (HierVL-w/o Summ) fails on long-term tasks (20.0% on SummaryMCQ); and (4) bidirectional parent loss LparentSN\mathcal{L}_{\text{parent}}^{SN} (Summ $\leftrightarrow$ Narr) provides consistent gains across tasks.

  6. Knowl 6 — Ego4D Long-Term Action Anticipation Performance

    data/table

    The Ego4D Long Term Anticipation (LTA) challenge requires predicting the next Z=20Z = 20 future actions (verb and noun sequences) given the current action. Performance is measured using Edit Distance (ED, lower is better) on the Ego4D test set.

    Method Verb ED ↓\downarrow Noun ED ↓\downarrow Act. ED ↓\downarrow
    Ego4D baseline 0.7389 0.7800 0.9432
    Robovision 0.7389 0.7688 0.9412
    I-CVAE 0.7526 0.7489 0.9308
    HierVL-w/o Hier 0.7691 0.7454 0.9451
    HierVL-Avg (Ours) 0.7223 0.7527 0.9401
    HierVL-SA (Ours) 0.7239 0.7349 0.9275

    HierVL-SA achieves the best results across all evaluated methods with a Verb ED of 0.7239, Noun ED of 0.7349, and Action ED of 0.9275. The non-hierarchical baseline HierVL-w/o Hier performs substantially worse (0.9451 Action ED) despite having access to the same video summaries, demonstrating that hierarchical feature aggregation is necessary for forecasting extended action sequences.

  7. Knowl 7 — Charades-Ego Action Recognition Transfer Results

    data/table

    Evaluations on Charades-Ego action recognition across 157 categories measure transferability in both zero-shot and fine-tuned settings using mean Average Precision (mAP).

    Zero-shot
    Method Task ckpt mAP PT ckpt mAP
    EgoVLP 25.0 19.4
    HierVL-w/o Hier 24.6 24.5
    HierVL-Avg (Ours) 25.2 23.9
    HierVL-SA (Ours) 26.0 25.0
    Fine-tuned
    Method mAP
    Actor 20.0
    SSDA 23.1
    I3D 25.8
    Ego-Exo 30.1
    EgoVLP 32.1
    HierVL-w/o Hier 32.6
    HierVL-Avg (Ours) 32.6
    HierVL-SA (Ours) 33.8

    In the zero-shot setting, EgoVLP suffers from severe overfitting on Ego4D, exhibiting a drop from 25.0% mAP (task-selected checkpoint) to 19.4% mAP (best pretraining checkpoint). In contrast, HierVL-SA resists pretraining overfitting, achieving 25.0% mAP at the best pretraining checkpoint and 26.0% mAP at the task checkpoint. When fine-tuned on Charades-Ego, HierVL-SA achieves 33.8% mAP, setting the state-of-the-art and outperforming EgoVLP by 1.7% mAP.

  8. Knowl 8 — EPIC-KITCHENS-100 Multi-Instance Retrieval Evaluation

    data/table

    EPIC-KITCHENS-100 Multi-Instance Retrieval (MIR) measures cross-modal retrieval across video-to-text (V→TV \to T) and text-to-video (T→VT \to V) using mean Average Precision (mAP) and normalized Discounted Cumulative Gain (nDCG), averaged over both directions.

    Zero-shot
    Method mAP Avg nDCG Avg
    EgoVLP 16.6 23.1
    HierVL-w/o Hier 17.8 24.1
    HierVL-Avg (Ours) 16.7 23.5
    HierVL-SA (Ours) 18.9 24.7
    Fine-tuned
    Method mAP Avg nDCG Avg
    MI-MM w/ S3D 29.2 44.7
    MME w/ TBN 38.5 48.5
    JPoSE w/ TBN 44.0 53.5
    EgoVLP 45.0 59.4
    HierVL-w/o Hier 44.7 59.8
    HierVL-Avg (Ours) 44.9 59.8
    HierVL-SA (Ours) 46.7 61.1

    In the zero-shot setting (where fvf_v and fnf_n are frozen), HierVL-SA achieves 18.9% mAP and 24.7% nDCG, outperforming EgoVLP by 2.3% mAP and 1.6% nDCG. In the fine-tuned setting (where fvf_v and fnf_n are updated), HierVL-SA attains 46.7% mAP and 61.1% nDCG, surpassing the prior state-of-the-art EgoVLP (45.0% mAP / 59.4% nDCG).

  9. Knowl 9 — Linear Probing Classification on HowTo100M Long Videos

    data/table

    To evaluate representation transfer on third-person instructional video without fine-tuning feature extractors, linear probing was conducted on the 100 most frequent classes in HowTo100M. The encoders fv,fnf_v, f_n and aggregator Agg\text{Agg} are frozen, and a single linear layer (25.7K parameters) is trained on top of the aggregated video representations.

    Method Inference Aggregator Classification Accuracy (%)
    EgoVLP Average 53.4
    HierVL-SA (Ours) Self-attention 54.6
    HierVL-Avg (Ours) Average 63.3
    HierVL-SA (Ours) Average 64.6

    All HierVL variants significantly outperform the EgoVLP baseline (53.4%). While self-attention inference achieves 54.6%, parameter-free average pooling at inference yields 63.3% for HierVL-Avg and 64.6% for HierVL-SA. The 1.3% gain of HierVL-SA over HierVL-Avg under identical average pooling inference demonstrates that the clip-level visual encoder fvf_v learned via HierVL-SA is inherently more expressive.

  10. Knowl 10 — HierVL Architectural and Implementation Specifications

    experimental setup

    HierVL employs the following architecture and pretraining configuration on the Ego4D dataset (comprising 3.8M narrations and 120K long-term summaries across 3,670 hours of video):

    • Visual Encoder (fvf_v): FrozenInTime video backbone based on TimeSformer and Vision Transformer (ViT), learned from scratch. Video clips are sampled at 1 fps, and the output representation is taken from the final [CLS] token.
    • Text Encoder (fnf_n): DistilBERT architecture used for both short-term action narrations nn and long-term activity summaries SS.
    • Aggregator (Agg\text{Agg}): 6-layer self-attention block derived from TimeSformer for HierVL-SA; parameter-free feature averaging for HierVL-Avg. In both variants, K=16K = 16 short-term feature representations are uniformly sampled across the video.
    • Optimization and Hardware: Pretrained across 4 nodes (32 total NVIDIA V100 32 GB GPUs) for 10 epochs over 2 days using the AdamW optimizer with a learning rate of 3×10−53 \times 10^{-5}. Batch sizes are 16 per GPU for short-term contrastive learning and 1 per GPU (consisting of 16 clips from the same video) for long-term video-level aggregation training, executed every m=5m = 5 epochs.

Coverage note — None was omitted; all primary architectural components, mathematical formulations, training algorithms, ablation variants, and downstream empirical benchmark results have been captured in the knowls.

References

  1. 1.Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. Youtube-8m: A large-scale video classification benchmark. arXiv preprint arXiv:1609.08675, 2016.
  2. 2.Yazan Abu Farha, Alexander Richard, and Juergen Gall. When will you do what?-anticipating temporal occurrences of activities. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5343–5352, 2018.
  3. 3.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1728–1738, 2021.
  4. 4.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. A clip-hitchhiker's guide to long video retrieval. arXiv preprint arXiv:2205.08508, 2022.
  5. 5.Siddhant Bansal, Chetan Arora, and CV Jawahar. My view is the best view: Procedure learning from egocentric videos. arXiv preprint arXiv:2207.10883, 2022.
  6. 6.Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020.
  7. 7.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In ICML, volume 2, page 4, 2021.
  8. 8.Jing Bi, Jiebo Luo, and Chenliang Xu. Procedure planning in instructional videos via contextual modeling and model-based policy learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15611–15620, 2021.
  9. 9.Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition, pages 961–970, 2015.
  10. 10.Chien-Yi Chang, De-An Huang, Danfei Xu, Ehsan Adeli, Li Fei-Fei, and Juan Carlos Niebles. Procedure planning in instructional videos. In European Conference on Computer Vision, pages 334–350. Springer, 2020.
  11. 11.Xing Cheng, Hezheng Lin, Xiangyu Wu, Fan Yang, and Dong Shen. Improving video-text retrieval by multi-stream corpus alignment and dual softmax loss. arXiv preprint arXiv:2109.04290, 2021.
  12. 12.Jinwoo Choi, Gaurav Sharma, Manmohan Chandraker, and Jia-Bin Huang. Unsupervised and semi-supervised domain adaptation for action recognition from drones. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1717–1726, 2020.
  13. 13.Ego4D Consortium. Egocentric live 4d perception (Ego4D) database: A large-scale first-person video database, supporting research in multi-modal machine perception for daily life activity. https://sites.google.com/view/ego4d/home.
  14. 14.Bo Dai, Yuqi Zhang, and Dahua Lin. Detecting visual relationships with deep relational networks. In Proceedings of the IEEE conference on computer vision and Pattern recognition, pages 3076–3086, 2017.
  15. 15.Dima Damen, Hazel Doughty, Giovanni Maria Farinella, , Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International Journal of Computer Vision (IJCV), 130:33–55, 2022.
  16. 16.Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Rescaling egocentric vision: collection, pipeline and challenges for epic-kitchens-100. International Journal of Computer Vision, 130(1):33–55, 2022.
  17. 17.Srijan Das and Michael S Ryoo. Video+ clip baseline for ego4d long-term action anticipation. arXiv preprint arXiv:2207.00579, 2022.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  19. 19.Haiwen Diao, Ying Zhang, Lin Ma, and Huchuan Lu. Similarity reasoning and filtration for image-text matching. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1218–1226, 2021.
  20. 20.Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2625–2634, 2015.
  21. 21.Qi Dong, Zhuowen Tu, Haofu Liao, Yuting Zhang, Vijay Mahadevan, and Stefano Soatto. Visual relationship detection using part-and-sum transformers with composite queries. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3550–3559, 2021.
  22. 22.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  23. 23.Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen. Clip2video: Mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097, 2021.
  24. 24.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6202–6211, 2019.
  25. 25.Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, and Zicheng Liu. Violet: End-to-end video-language transformers with masked visual-token modeling. arXiv preprint arXiv:2111.12681, 2021.
  26. 26.Antonino Furnari and Giovanni Maria Farinella. Rolling-unrolling lstms for action anticipation from first-person video. IEEE transactions on pattern analysis and machine intelligence, 43(11):4021–4036, 2020.
  27. 27.Adrien Gaidon, Zaid Harchaoui, and Cordelia Schmid. Temporal localization of actions with actoms. IEEE transactions on pattern analysis and machine intelligence, 35(11):2782–2795, 2013.
  28. 28.Jiyang Gao, Zhenheng Yang, and Ram Nevatia. Red: Reinforced encoder-decoder networks for action anticipation. arXiv preprint arXiv:1707.04818, 2017.
  29. 29.Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, and Thomas Brox. Coot: Cooperative hierarchical transformer for video-text representation learning. Advances in neural information processing systems, 33:22605–22618, 2020.
  30. 30.Rohit Girdhar and Kristen Grauman. Anticipative video transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13505–13515, 2021.
  31. 31.Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell. Actionvlad: Learning spatio-temporal aggregation for action classification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 971–980, 2017.
  32. 32.Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra. Omnivore: A single model for many visual modalities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16102–16112, 2022.
  33. 33.Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio. An empirical investigation of catastrophic forgetting in gradient-based neural networks. arXiv preprint arXiv:1312.6211, 2013.
  34. 34.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18995–19012, 2022.
  35. 35.Albert Gu, Karan Goel, and Christopher Ré. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396, 2021.
  36. 36.Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang. Scaling up vision-language pre-training for image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17980–17989, 2022.
  37. 37.Yan Huang, Qi Wu, Chunfeng Song, and Liang Wang. Learning semantic concepts and order for image and sentence matching. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6163–6171, 2018.
  38. 38.Noureldien Hussein, Efstratios Gavves, and Arnold WM Smeulders. Timeception for complex action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 254–263, 2019.
  39. 39.Md Mohaiminul Islam and Gedas Bertasius. Long movie clip classification with state-space video models. arXiv preprint arXiv:2204.01692, 2022.
  40. 40.Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen. Epic-fusion: Audio-visual temporal binding for egocentric action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5492–5501, 2019.
  41. 41.Ronald Kemker, Marc McClure, Angelina Abitino, Tyler Hayes, and Christopher Kanan. Measuring catastrophic forgetting in neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018.
  42. 42.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526, 2017.
  43. 43.Bruno Korbar, Du Tran, and Lorenzo Torresani. Scsampler: Sampling salient clips from video for efficient action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6232–6242, 2019.
  44. 44.Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. Stacked cross attention for image-text matching. In Proceedings of the European conference on computer vision (ECCV), pages 201–216, 2018.
  45. 45.Jie Lei, Tamara L Berg, and Mohit Bansal. Revealing single frame bias for video-and-language learning. arXiv preprint arXiv:2206.03428, 2022.
  46. 46.Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. Less is more: Clipbert for video-and-language learning via sparse sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7331–7341, 2021.
  47. 47.Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal. Tvqa+: Spatio-temporal grounding for video question answering. arXiv preprint arXiv:1904.11574, 2019.
  48. 48.Kunchang Li, Yali Wang, Junhao Zhang, Peng Gao, Guanglu Song, Yu Liu, Hongsheng Li, and Yu Qiao. Uniformer: Unifying convolution and self-attention for visual recognition. arXiv preprint arXiv:2201.09450, 2022.
  49. 49.Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu. Visual semantic reasoning for image-text matching. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4654–4662, 2019.
  50. 50.Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. Hero: Hierarchical encoder for video+ language omni-representation pre-training. arXiv preprint arXiv:2005.00200, 2020.
  51. 51.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision, pages 121–137. Springer, 2020.
  52. 52.Yanghao Li, Tushar Nagarajan, Bo Xiong, and Kristen Grauman. Ego-exo: Transferring visual representations from third-person to first-person videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6943–6953, 2021.
  53. 53.Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Mvitv2: Improved multiscale vision transformers for classification and detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4804–4814, 2022.
  54. 54.Ji Lin, Chuang Gan, and Song Han. Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7083–7093, 2019.
  55. 55.Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, et al. Egocentric video-language pretraining. In NeurIPS, 2022.
  56. 56.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  57. 57.Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao, Chenglin Wu, Cheng Deng, and Rongrong Ji. Multi-task collaborative network for joint referring expression comprehension and segmentation. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pages 10034–10043, 2020.
  58. 58.Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Jason Li, Taroon Bharti, and Ming Zhou. Univl: A unified video and language pre-training model for multimodal understanding and generation. arXiv preprint arXiv:2002.06353, 2020.
  59. 59.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 11–20, 2016.
  60. 60.Esteve Valls Mascaro, Hyemin Ahn, and Dongheui Lee. Intention-conditioned long-term human egocentric action forecasting@ ego4d challenge 2022. arXiv preprint arXiv:2207.12080, 2022.
  61. 61.Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. End-to-end learning of visual representations from uncurated instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9879–9889, 2020.
  62. 62.Antoine Miech, Ivan Laptev, and Josef Sivic. Learnable pooling with context gating for video classification. arXiv preprint arXiv:1706.06905, 2017.
  63. 63.Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2630–2640, 2019.
  64. 64.Zwe Naing and Ehsan Elhamifar. Procedure completion by learning from partial summaries. In British Machine Vision Conference, 2020.
  65. 65.Van-Quang Nguyen, Masanori Suganuma, and Takayuki Okatani. Grit: Faster and better image captioning transformer using dual visual features. arXiv preprint arXiv:2207.09666, 2022.
  66. 66.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  67. 67.Hamed Pirsiavash and Deva Ramanan. Parsing videos of actions with segmental grammars. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 612–619, 2014.
  68. 68.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR, 18–24 Jul 2021.
  69. 69.Santhosh Ramakrishnan, Ziad Al-Halah, and Kristen Grauman. Naq: Leveraging narrations as queries to supervise episodic memory. In CVPR, 2023.
  70. 70.Adrià Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Florian Strub, Corentin Tallec, Mateusz Malinowski, Viorica Pătrăucean, Florent Altché, Michal Valko, et al. Broaden your views for self-supervised video learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1255–1265, 2021.
  71. 71.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019.
  72. 72.Hengcan Shi, Hongliang Li, Fanman Meng, and Qingbo Wu. Key-word-aware network for referring expression image segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), pages 38–54, 2018.
  73. 73.Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari. Actor and observer: Joint modeling of first and third-person videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7396–7404, 2018.
  74. 74.Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari. Charades-ego: A large-scale dataset of paired third and first person videos. arXiv preprint arXiv:1804.09626, 2018.
  75. 75.Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216, 2019.
  76. 76.Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. Movieqa: Understanding stories in movies through question-answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4631–4640, 2016.
  77. 77.Gül Varol, Ivan Laptev, and Cordelia Schmid. Long-term temporal convolutions for action recognition. IEEE transactions on pattern analysis and machine intelligence, 40(6):1510–1517, 2017.
  78. 78.Jue Wang, Gedas Bertasius, Du Tran, and Lorenzo Torresani. Long-short temporal contrastive learning of video transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14010–14020, June 2022.
  79. 79.Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Luowei Zhou, Yucheng Zhao, Yujia Xie, Ce Liu, Yu-Gang Jiang, and Lu Yuan. Omnivl: One foundation model for image-language and video-language tasks. arXiv preprint arXiv:2209.07526, 2022.
  80. 80.Jue Wang and Anoop Cherian. Learning discriminative video representations using adversarial perturbations. In Proceedings of the European Conference on Computer Vision (ECCV), pages 685–701, 2018.
  81. 81.Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang. Git: A generative image-to-text transformer for vision and language. arXiv preprint arXiv:2205.14100, 2022.
  82. 82.Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. Temporal segment networks: Towards good practices for deep action recognition. In European conference on computer vision, pages 20–36. Springer, 2016.
  83. 83.Michael Wray, Diane Larlus, Gabriela Csurka, and Dima Damen. Fine-grained action retrieval through multiple parts-of-speech embeddings. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 450–459, 2019.
  84. 84.Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick. Long-term feature banks for detailed video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 284–293, 2019.
  85. 85.Chao-Yuan Wu and Philipp Krahenbuhl. Towards long-form video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1884–1894, 2021.
  86. 86.Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13587–13597, 2022.
  87. 87.Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu, R Manmatha, Alexander J Smola, and Philipp Krähenbühl. Compressed video action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6026–6035, 2018.
  88. 88.Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In Proceedings of the European conference on computer vision (ECCV), pages 305–321, 2018.
  89. 89.Hu Xu, Gargi Ghosh, Po-Yao Huang, Prahal Arora, Masoumeh Aminzadeh, Christoph Feichtenhofer, Florian Metze, and Luke Zettlemoyer. Vlm: Task-agnostic video-language model pre-training for video understanding. arXiv preprint arXiv:2105.09996, 2021.
  90. 90.Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. Videoclip: Contrastive pre-training for zero-shot video-text understanding. arXiv preprint arXiv:2109.14084, 2021.
  91. 91.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296, 2016.
  92. 92.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Just ask: Learning to answer questions from millions of narrated videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1686–1697, 2021.
  93. 93.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Zero-shot video question answering via frozen bidirectional language models. arXiv preprint arXiv:2206.08155, 2022.
  94. 94.Jianwei Yang, Yonatan Bisk, and Jianfeng Gao. Taco: Token-aware cascade contrastive learning for video-text alignment. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11562–11572, 2021.
  95. 95.Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. Beyond short snippets: Deep networks for video classification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4694–4702, 2015.
  96. 96.Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. Beyond short snippets: Deep networks for video classification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4694–4702, 2015.
  97. 97.Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. Merlot: Multimodal neural script knowledge models. Advances in Neural Information Processing Systems, 34:23634–23651, 2021.
  98. 98.Yan Zeng, Xinsong Zhang, and Hang Li. Multi-grained vision language pre-training: Aligning texts with visual concepts. arXiv preprint arXiv:2111.08276, 2021.
  99. 99.Yan Zeng, Xinsong Zhang, and Hang Li. Multi-grained vision language pre-training: Aligning texts with visual concepts. arXiv preprint arXiv:2111.08276, 2021.
  100. 100.Bowen Zhang, Hexiang Hu, and Fei Sha. Cross-modal and hierarchical modeling of video and text. In Proceedings of the european conference on computer vision (ECCV), pages 374–390, 2018.
  101. 101.Chenlin Zhang, Jianxin Wu, and Yin Li. Actionformer: Localizing moments of actions with transformers. arXiv preprint arXiv:2202.07925, 2022.
  102. 102.He Zhao, Isma Hadji, Nikita Dvornik, Konstantinos G Derpanis, Richard P Wildes, and Allan D Jepson. P3iv: Probabilistic procedure planning from instructional videos with weak supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2938–2948, 2022.
  103. 103.Shuai Zhao, Linchao Zhu, Xiaohan Wang, and Yi Yang. Centerclip: Token clustering for efficient text-video retrieval. arXiv preprint arXiv:2205.00823, 2022.
  104. 104.Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba. Temporal relational reasoning in videos. In Proceedings of the European conference on computer vision (ECCV), pages 803–818, 2018.
  105. 105.Luowei Zhou, Chenliang Xu, and Jason J Corso. Towards automatic learning of procedures from web instructional videos. In AAAI Conference on Artificial Intelligence, pages 7590–7598, 2018.
  106. 106.Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic. Cross-task weakly supervised learning from instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3537–3545, 2019.
  107. 107.Mohammadreza Zolfaghari, Kamaljeet Singh, and Thomas Brox. Eco: Efficient convolutional network for online video understanding. In Proceedings of the European conference on computer vision (ECCV), pages 695–712, 2018.

Citation

MLA
Ashutosh, K., et al. “HierVL: Learning Hierarchical Video-Language Embeddings”. arXiv, 2023, http://arxiv.org/abs/2301.02311v2.
APA
Ashutosh, K., Girdhar, R., Torresani, L., & Grauman, K. (2023). HierVL: Learning Hierarchical Video-Language Embeddings. arXiv. http://arxiv.org/abs/2301.02311v2
Chicago
Ashutosh, K., R. Girdhar, L. Torresani, and K. Grauman. 2023. “HierVL: Learning Hierarchical Video-Language Embeddings”. arXiv. http://arxiv.org/abs/2301.02311v2.
Harvard
Ashutosh, K. et al. (2023) “HierVL: Learning Hierarchical Video-Language Embeddings”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.02311v2.
Vancouver
1. Ashutosh K, Girdhar R, Torresani L, Grauman K (2023) HierVL: Learning Hierarchical Video-Language Embeddings. arXiv

BibTeX

@article{ashutosh2023hiervl,
  title = {HierVL: Learning Hierarchical Video-Language Embeddings},
  author = {Ashutosh, Kumar and Girdhar, Rohit and Torresani, Lorenzo and Grauman, Kristen},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.02311v2},
  eprint = {2301.02311}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE