Video ReCap: Recursive Captioning of Hour-Long Videos

Md Mohaiminul IslamNgan HoXitong YangTushar NagarajanLorenzo TorresaniGedas Bertasius

article2024CVPR109 citations

Presents a recursive video-language model and benchmark dataset that efficiently generate hierarchical captions across multiple temporal granularities for hour-long untrimmed videos.

Listen

Real-world video understanding poses a significant challenge because most natural videos last for minutes or hours and contain complex, nested human behaviors. However, current automated video captioning systems are primarily built for short clips spanning only 5 to 15 seconds. When applied to extended durations, existing models become computationally prohibitive, fail to compress repetitive visual details, and cannot connect immediate actions to overarching goals. As long-form video archives grow rapidly across industries, automated tools must be able to understand visual content across multiple time scales without requiring unsustainable computing resources.

The article introduces and evaluates Video ReCap, a recursive video-language model designed to generate natural language captions across multiple hierarchical levels for videos ranging from several seconds to two hours in duration. The article also introduces Ego4D-HCap, a newly curated benchmark designed to evaluate multi-level captioning on long, untrimmed egocentric video.

The researchers developed a recursive framework that processes long videos in three distinct stages: short clip-level captions describing atomic actions, medium-length segment descriptions covering intermediate steps, and long-range video summaries capturing high-level goals. The architecture uses an off-the-shelf visual encoder, a video-language alignment module built on frozen language models that compresses large volumes of features into compact representations, and a recursive text decoder. Captions generated at lower tiers serve as text inputs for higher tiers alongside sparsely sampled video features. To train the system effectively despite severe data scarcity at longer timescales, the authors implemented a psychologically inspired curriculum learning strategy—moving sequentially from short clips to full summaries—and used large language models to generate pseudo-annotations to expand training data.

The findings demonstrate substantial performance gains across all temporal granularities. Video ReCap outperformed leading baselines, including specialized models such as LaViLa and multimodal zero-shot approaches. For long-range video summaries, Video ReCap achieved a CIDEr metric score of 28.06 compared to 6.54 for standard LaViLa and 20.12 for heavily parameterized language model baselines. Incorporating large language model supervision further raised the segment description score to 46.88 and the full summary score to 29.34. The unified variant, Video ReCap-U, achieved highly competitive performance while utilizing only 113 million trainable parameters compared to 258 million to 586 million in comparative systems. Finally, applying the model’s generated hierarchical captions to long-form video question answering on the EgoSchema benchmark established a new state-of-the-art accuracy of 50.23%, outperforming the prior benchmark leader by 18.13 percentage points.

These results demonstrate that long-range video understanding does not require ingesting complete, uncompressed video streams at massive computational expense. Instead, passing concise text summaries and sparse visual features upward through a hierarchical pipeline enables scalable reasoning over hour-long footage. This recursive approach significantly lowers the operational and memory costs required for long-form video analysis while improving comprehension of intent, intermediate procedures, and complex context.

Organizations handling extensive video data should consider adopting hierarchical processing workflows rather than relying on flat, clip-by-clip analyses or brute-force vision models. Technical teams should explore the unified model architecture (Video ReCap-U) when computing budgets are constrained, as it delivers high accuracy with a substantially smaller parameter footprint. Future development efforts should focus on extending these recursive methods toward real-time stream captioning, conversational video interfaces, and interactive query systems.

The study's primary limitation lies in its primary evaluation on egocentric, activity-focused datasets (Ego4D), which may exhibit different narrative and temporal structures compared to third-person broadcast media, surveillance, or specialized industrial feeds. Additionally, generating higher-tier summaries relies partially on the quality of preceding lower-tier captions, creating potential risks of cascading errors if early action recognition fails. Nonetheless, the consistent performance gains across both captioning and complex question-answering benchmarks provide high confidence in the fundamental recursive methodology.

  • Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). This technical report advances ultra-long video context understanding and fine-grained visual-language reasoning using explicit textual timestamps and deep cross-layer token injection.
Cover for Video ReCap: Recursive Captioning of Hour-Long Videos

Abstract

Most video captioning models are designed to process short video clips of few seconds and output text describing low-level visual concepts (e.g., objects, scenes, atomic actions). However, most real-world videos last for minutes or hours and have a complex hierarchical structure spanning different temporal granularities. We propose Video ReCap, a recursive video captioning model that can process video inputs of dramatically different lengths (from 1 second to 2 hours) and output video captions at multiple hierarchy levels. The recursive video-language architecture exploits the synergy between different video hierarchies and can process hour-long videos efficiently. We utilize a curriculum learning training scheme to learn the hierarchical structure of videos, starting from clip-level captions describing atomic actions, then focusing on segment-level descriptions, and concluding with generating summaries for hour-long videos. Furthermore, we introduce Ego4D-HCap dataset by augmenting Ego4D with 8,267 manually collected long-range video summaries. Our recursive model can flexibly generate captions at different hierarchy levels while also being useful for other complex video understanding tasks, such as VideoQA on EgoSchema. Data, code, and models are publicly available at https://sites.google.com/view/vidrecap.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Technical Approach
  • 3.1. Problem Overview
  • 3.2. Recursive Video-Language Model
  • 3.3. Hierarchical Curriculum Learning
  • 3.4. Additional Supervision using Language Models
  • 3.5. Implementation Details
  • 4. Ego4D-HCap Dataset
  • 5. Experimental Setup
  • 5.1. Hierarchical Video Captioning Baselines
  • 5.2. Our Model Variants
  • 6. Results and Analysis
  • 6.1. Hierarchical Video Captioning Results
  • 6.2. Long-Range VideoQA on EgoSchema
  • 6.3. Ablation Studies
  • 7. Conclusions and Future Work
  • References

Knowls

  1. Knowl 1 — Hierarchical Video Captioning Task Formulation

    equation

    Given an untrimmed long-range video sequence Vi=[Ii(t)]t=1TV_i = [I_i^{(t)}]_{t=1}^T comprising TT RGB frames, hierarchical video captioning seeks to generate textual descriptions across three distinct temporal hierarchies ℓ∈{1,2,3}\ell \in \{1, 2, 3\}:

    1. Clip Captions (Yi(1)Y_i^{(1)}): Short-range descriptions of fine-grained atomic actions and objects occurring within intervals of several seconds.
    2. Segment Descriptions (Yi(2)Y_i^{(2)}): Intermediate-range descriptions capturing multi-step activities or scenes unfolding over several minutes.
    3. Video Summaries (Yi(3)Y_i^{(3)}): Long-range summaries characterizing overall intent, character interactions, and overarching activities across entire videos lasting up to several hours.

    Each caption sequence Yi(ℓ)=[yi,j(ℓ)]j=1∣Yi(ℓ)∣Y_i^{(\ell)} = [y_{i,j}^{(\ell)}]_{j=1}^{|Y_i^{(\ell)}|} is generated autoregressively conditioned on visual features XiX_i extracted from the video and the caption sequence Yi(ℓ−1)Y_i^{(\ell-1)} produced at the preceding hierarchy level according to the likelihood objective:

    p(Y(ℓ)∣X)=∏k=1Kp(yk(ℓ)∣y<k(ℓ),X,Y(ℓ−1))p(Y^{(\ell)} \mid X) = \prod_{k=1}^K p(y_k^{(\ell)} \mid y_{< k}^{(\ell)}, X, Y^{(\ell-1)})

    where yk(ℓ)y_k^{(\ell)} denotes the kk-th text token of the target caption at hierarchy level ℓ\ell, y<k(ℓ)y_{< k}^{(\ell)} represents preceding tokens in that caption, and Y(0)=∅Y^{(0)} = \emptyset serves as the recursive base case.

  2. Knowl 2 — Video ReCap Recursive Video-Language Architecture

    model/method

    The Video ReCap model generates hierarchical captions via three sequential components:

    1. Video Encoder: An off-the-shelf video Transformer (TimeSformer) extracts features Xi=[xi,j]j=1∣C∣X_i = [x_{i,j}]_{j=1}^{|C|} across ∣C∣|C| uniform video clips. For short-clip captions (level ℓ=1\ell=1), dense spatiotemporal feature tensors x∈RF×H×W×Dx \in \mathbb{R}^{F \times H \times W \times D} (where FF is frames per clip, H×WH \times W is spatial resolution, and DD is feature dimensionality) are retained. For segment descriptions (ℓ=2\ell=2) and full video summaries (ℓ=3\ell=3), sparsely sampled global CLS tokens are extracted instead to constrain compute over long durations.
    2. Video-Language (VL) Alignment Module: Compresses video features XiX_i and preceding-level generated captions Yi(ℓ−1)Y_i^{(\ell-1)} into a joint sequence of ∣Z∣|Z| embeddings Zi=[zi,j]j=1∣Z∣Z_i = [z_{i,j}]_{j=1}^{|Z|} with dimension DzD_z (e.g., ∣Z∣=256|Z|=256). It uses a frozen language model (DistilBERT) with injected trainable cross-attention layers in each block to encode video features XiX_i into fixed video tokens, alongside an identical frozen LM branch to encode prior captions Yi(ℓ−1)Y_i^{(\ell-1)} into fixed text tokens. The two sets of tokens are concatenated to form ZiZ_i. For the clip-caption base case (ℓ=1\ell=1), ZiZ_i consists entirely of video tokens.
    3. Recursive Text Decoder: A pretrained language model (GPT-2, 12 layers, hidden dimension 768) augmented with trainable cross-attention layers in each transformer block that attend to the joint embeddings ZiZ_i.

    Two structural configurations are used:

    • Video ReCap: Employs distinct, non-shared decoders and alignment modules for each hierarchy level (339M total trainable parameters).
    • Video ReCap-U: Shares a single set of alignment and decoder parameters across all hierarchy levels (113M trainable parameters).
  3. Knowl 3 — Hierarchical Curriculum Learning Training Scheme

    model/method

    To manage severe data imbalance across temporal tiers and learn dependencies from atomic actions to high-level goals, Video ReCap is trained using a progressive curriculum learning strategy:

    1. Stage 1 (Clip Captioning): The model is initialized with pretrained weights and trained exclusively on atomic action descriptions using dense spatiotemporal clip features.
    2. Stage 2 (Segment Descriptions): The model weights from Stage 1 are loaded and finetuned on medium-length video segments using sparse visual features concatenated with Stage 1 generated clip descriptions.
    3. Stage 3 (Video Summaries): The model weights from Stage 2 are loaded and finetuned on whole-video summaries using sparse visual features concatenated with Stage 2 generated segment descriptions.

    This progression allows lower-level perceptual features to stabilize before conditioning higher-level semantic summarization on lower-level generated text.

  4. Knowl 4 — LLM Teacher Pseudo-Supervision for Long-Range Summaries

    model/method

    To alleviate the scarcity of human annotations for long-range egocentric video segments and full videos, an LLM-based pseudo-annotation pipeline provides auxiliary training data:

    1. Teacher Fine-Tuning: A sequence-to-sequence language model (FLAN-T5-Large) is finetuned on available ground-truth hierarchical text to map sequences of short-term clip captions concatenated across varying time windows into cohesive medium-length segment descriptions and whole-video summaries.
    2. Pseudo-Data Generation: The trained teacher generates 100,000 synthetic segment descriptions and 15,000 synthetic video summaries from concatenated ground-truth clip captions.
    3. Augmented Training: Video ReCap is trained on the union of manual human annotations and LLM-generated pseudo-annotations.
  5. Knowl 5 — Ego4D-HCap Dataset Benchmark

    data/table

    Ego4D-HCap is a hierarchical video captioning benchmark constructed on top of the Ego4D egocentric video dataset by annotating 8,267 untrimmed long-form videos with human-written summaries for durations spanning up to two hours.

    Hierarchy Level Number of Samples Average Duration
    Clip Caption 5.27M 0.96 sec
    Segment Description 17.5K 2.87 min
    Video Summary 8.3K 28.46 min

    The resulting dataset spans three temporal tiers: time-stamped atomic actions (clips), intermediate activity step descriptions spanning several minutes (segments), and full-length overarching goal summaries (video summaries).

  6. Knowl 6 — Hierarchical Video Captioning Evaluation on Ego4D-HCap

    data/table

    Video ReCap variants were evaluated against zero-shot and fully-finetuned video-language baselines on the Ego4D-HCap test split using CIDEr (C), ROUGE-L (R), and METEOR (M).

    Model Clip Caption Segment Description Video Summary
    C R M C R M C R
    BLIP2 (Zero-Shot) 8.10 7.40 12.70 - - - - -
    BLIP2 + GPT3.5 - - - 5.68 16.87 13.47 11.13 22.41
    LaViLa + GPT3.5 - - - 5.79 19.77 13.45 12.16 24.49
    LaViLa + GPT2 - - - 38.22 38.10 16.58 17.98 29.48
    LaViLa + FLAN-T5 - - - 39.13 38.77 16.88 20.12 30.06
    LaViLa (End-to-End) 88.56 47.64 28.03 24.63 33.31 15.30 6.54 23.97
    Video ReCap (w/o Pseudo-Ann.) 98.35 48.77 28.28 41.74 39.04 18.21 28.06 32.27
    Video ReCap (Full) 98.35 48.77 28.28 46.88 39.73 18.55 29.34 32.64
    Video ReCap-U (Unified) 92.67 47.90 28.08 45.60 39.33 18.17 31.06 33.32

    Video ReCap outperforms LaViLa by 9.79 CIDEr on clip captioning, 22.25 CIDEr on segment descriptions, and 22.80 CIDEr on full video summaries. The unified variant Video ReCap-U achieves the highest CIDEr on video summaries (31.06) using 113M trainable parameters compared to 586M for LaViLa + FLAN-T5 (20.12 CIDEr).

  7. Knowl 7 — Zero-Shot Long-Form VideoQA on EgoSchema via Hierarchical Captions

    data/table

    Evaluation on the EgoSchema benchmark (over 5,000 human-curated multiple-choice questions over 250 hours of video) tested zero-shot question answering using GPT-3.5 prompted with generated captions. All EgoSchema-overlapping videos were removed from Ego4D pretraining.

    Model Input Feature Ego4D Pretrained QA Acc (%)
    Random Choice - No 20.00
    GPT-3.5 Question Only No 19.57
    FrozenBiLM Video No 26.90
    VIOLET Video No 19.90
    mPLUG-Owl Video No 31.10
    InternVideo Video No 32.10
    EgoVLP Video Yes 34.86
    EgoVLPv2 Video Yes 34.12
    LaViLa + GPT-3.5 Clip Captions Yes 44.27
    Video ReCap + GPT-3.5 Clip Captions Yes 46.03
    Video ReCap + GPT-3.5 Hierarchical Captions Yes 50.23

    Prompting GPT-3.5 with hierarchical captions (clips, segments, and full summaries) achieves 50.23% accuracy, outperforming direct video QA foundation models like InternVideo by 18.13% and EgoVLPv2 by 16.11%, while exceeding flat clip caption prompting by 4.20%.

  8. Knowl 8 — Multimodal Input Ablation: Visual Features versus Recursive Textual Representations

    data/table

    An ablation over input modalities fed to the alignment module for segment description and video summary tiers on Ego4D-HCap shows the impact of visual features versus recursive text tokens.

    Input Modality Segment Description Video Summary
    C R M C R M
    Video Only (Sparse CLS Features) 40.17 38.65 17.59 25.64 29.61 13.57
    Text Only (Recursive Lower-Level Text) 40.10 38.02 17.41 23.23 29.17 13.31
    Video + Text (Combined) 41.74 39.04 18.21 28.06 32.27 14.26

    Combining sparse visual features with recursive textual context from preceding hierarchy tiers improves CIDEr by +1.57 points over video-only on segment descriptions and by +2.42 points over video-only (+4.83 points over text-only) on full video summaries.

  9. Knowl 9 — Curriculum Learning Progression Ablation

    data/table

    An ablation on training initialization strategies demonstrates the necessity of stage-by-stage hierarchical curriculum learning on Ego4D-HCap (evaluated without pseudo-supervision).

    Training Progression Segment Description Video Summary
    C R M C R M
    Init →\rightarrow Segment 36.81 38.70 17.17 - - -
    Caption →\rightarrow Segment 41.74 39.04 18.21 - - -
    Init →\rightarrow Video - - - 8.62 26.33 11.24
    Caption →\rightarrow Video - - - 24.84 30.74 13.25
    Caption →\rightarrow Segment →\rightarrow Video - - - 28.06 32.27 14.26

    Training directly on summaries from standard initialization (extInit→Video ext{Init} \rightarrow \text{Video}) without curriculum results in an 8.62 CIDEr score. Progressing through all three stages (extCaption→Segment→Video ext{Caption} \rightarrow \text{Segment} \rightarrow \text{Video}) achieves 28.06 CIDEr, outperforming the two-stage shortcut (Caption→Video\text{Caption} \rightarrow \text{Video}) by 3.22 CIDEr points.

  10. Knowl 10 — LLM Teacher Model Selection and Supervision Ablation

    data/table

    Comparison of sequence-to-sequence language models for synthesizing pseudo-annotations from concatenated ground-truth clip captions, along with the performance gain when added to Video ReCap training.

    LLM Teacher Model Segment Description Video Summary
    C R M C R M
    GPT-2 96.47 46.96 23.13 40.06 33.06 14.76
    GPT-2-Large 104.30 47.68 23.15 43.18 33.86 15.00
    FLAN-T5-Small 95.61 46.16 22.30 43.27 34.19 14.69
    FLAN-T5-Large 125.67 50.61 26.06 52.08 36.99 19.93

    FLAN-T5-Large achieves the highest text generation metrics (125.67 CIDEr on segments, 52.08 CIDEr on summaries). Adding 100K segment pseudo-annotations and 15K video summary pseudo-annotations from FLAN-T5-Large to human ground-truth data improves Video ReCap from 41.74 to 46.88 CIDEr (+5.14) on segment descriptions and from 28.06 to 29.34 CIDEr (+1.28) on video summaries.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736, 2022. 4, 7
  2. 2.Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, and Kristen Grauman. Hiervl: Learning hierarchical video-language embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23066–23078, 2023. 3
  3. 3.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014. 1
  4. 4.Albert Bandura. Social cognitive theory: An agentic perspective. Asian journal of social psychology, 2(1):21–41, 1999. 1, 4
  5. 5.Satanjeev Banerjee and Alon Lavie. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72, 2005. 6
  6. 6.Siddhant Bansal, Chetan Arora, and CV Jawahar. My view is the best view: Procedure learning from egocentric videos. In European Conference on Computer Vision, pages 657–675. Springer, 2022. 3
  7. 7.Lorenzo Baraldi, Costantino Grana, and Rita Cucchiara. Hierarchical boundary-aware neural encoder for video captioning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1657–1666, 2017. 2
  8. 8.R.G. Barker and H.F. Wright. Midwest and Its Children: The Psychological Ecology of an American Town. Row, Peterson, 1954. 1, 4
  9. 9.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In ICML, page 4, 2021. 3, 5
  10. 10.Matthew Botvinick and David C Plaut. Doing without schema hierarchies: a recurrent connectionist approach to normal and impaired routine sequential action. Psychological review, 111(2):395, 2004. 1, 4
  11. 11.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. 5, 6, 7
  12. 12.David Chen and William B Dolan. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pages 190–200, 2011. 3
  13. 13.Sihan Chen, Xingjian He, Longteng Guo, Xinxin Zhu, Weining Wang, Jinhui Tang, and Jing Liu. Valor: Vision-audio-language omni-perception pretraining model and dataset. arXiv preprint arXiv:2304.08345, 2023. 1
  14. 14.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022. 6, 8
  15. 15.Richard P Cooper and Tim Shallice. Hierarchical schemas and goals in the control of sequential behavior. 2006. 1, 4
  16. 16.Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2625–2634, 2015. 2
  17. 17.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 2
  18. 18.Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, and Zicheng Liu. An empirical study of end-to-end video-language transformers with masked visual modeling. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22898–22909, 2022. 7
  19. 19.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18995–19012, 2022. 2, 5, 6
  20. 20.Chiori Hori, Takaaki Hori, Teng-Yok Lee, Ziming Zhang, Bret Harsham, John R Hershey, Tim K Marks, and Kazuhiko Sumi. Attention-based multimodal fusion for video description. In Proceedings of the IEEE international conference on computer vision, pages 4193–4202, 2017. 1, 2
  21. 21.Gabriel Huang, Bo Pang, Zhenhai Zhu, Clara Rivera, and Radu Soricut. Multimodal pretraining for dense video captioning. arXiv preprint arXiv:2011.11760, 2020. 3
  22. 22.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014. 5
  23. 23.Atsuhiro Kojima, Takeshi Tamura, and Kunio Fukunaga. Natural language description of human activities from video images based on concept hierarchy of actions. International Journal of Computer Vision, 50:171–184, 2002. 2
  24. 24.Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. Dense-captioning events in videos. In Proceedings of the IEEE international conference on computer vision, pages 706–715, 2017. 3
  25. 25.Weiyu Lan, Xirong Li, and Jianfeng Dong. Fluency-guided cross-lingual image captioning. In Proceedings of the 25th ACM international conference on Multimedia, pages 1549–1557, 2017. 2
  26. 26.Jie Lei, Liwei Wang, Yelong Shen, Dong Yu, Tamara L Berg, and Mohit Bansal. Mart: Memory-augmented recurrent transformer for coherent video paragraph captioning. arXiv preprint arXiv:2005.05402, 2020. 1, 2
  27. 27.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023. 2, 3, 5, 6
  28. 28.Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. Hero: Hierarchical encoder for video+language omni-representation pre-training. In Conference on Empirical Methods in Natural Language Processing, 2020. 3
  29. 29.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81, 2004. 6
  30. 30.Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Z XU, Difei Gao, Rong-Cheng Tu, Wenzhe Zhao, Weijie Kong, et al. Egocentric video-language pretraining. Advances in Neural Information Processing Systems, 35:7575–7586, 2022. 7
  31. 31.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2017. 5
  32. 32.Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Jason Li, Taroon Bharti, and Ming Zhou. Univl: A unified video and language pre-training model for multimodal understanding and generation. arXiv preprint arXiv:2002.06353, 2020. 1
  33. 33.Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A diagnostic benchmark for very long-form video language understanding. arXiv preprint arXiv:2308.09126, 2023. 2, 7
  34. 34.Pingbo Pan, Zhongwen Xu, Yi Yang, Fei Wu, and Yueting Zhuang. Hierarchical recurrent neural encoder for video representation with application to captioning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1029–1038, 2016. 2
  35. 35.Yingwei Pan, Ting Yao, Houqiang Li, and Tao Mei. Video captioning with transferred semantic attributes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6504–6512, 2017. 1, 2
  36. 36.Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai. Memory-attended recurrent network for video captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8347–8356, 2019. 1, 2
  37. 37.Shraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin, Hardik Shah, Mike Zheng Shou, Rama Chellappa, and Pengchuan Zhang. Egovlpv2: Egocentric video-language pre-training with fusion in the backbone. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5285–5297, 2023. 7
  38. 38.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019. 3, 5, 6, 8
  39. 39.Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele. Movie description. International Journal of Computer Vision, 123:94–120, 2017. 3
  40. 40.Marcus Rohrbach, Wei Qiu, Ivan Titov, Stefan Thater, Manfred Pinkal, and Bernt Schiele. Translating video content to natural language descriptions. In Proceedings of the IEEE international conference on computer vision, pages 433–440, 2013. 1, 2
  41. 41.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: Smaller, faster, cheaper and lighter. arxiv 2019. arXiv preprint arXiv:1910.01108, 2019. 3
  42. 42.Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21096–21106, 2022. 3
  43. 43.Paul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, and Cordelia Schmid. End-to-end generative pretraining for multimodal video captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17959–17968, 2022. 1, 2
  44. 44.Jingkuan Song, Yuyu Guo, Lianli Gao, Xuelong Li, Alan Hanjalic, and Heng Tao Shen. From deterministic to generative: Multimodal stochastic rnns for video captioning. IEEE transactions on neural networks and learning systems, 30 (10):3047–3058, 2018. 2
  45. 45.Yale Song, Gene Byrne, Tushar Nagarajan, Huiyu Wang, Miguel Martin, and Lorenzo Torresani. Ego4d goal-step: Toward hierarchical understanding of procedural activities. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. 3
  46. 46.Chen Sun and Ram Nevatia. Semantic aware video transcription using random forest classifiers. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part I 13, pages 772–786. Springer, 2014. 2
  47. 47.Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. Videobert: A joint model for video and language representation learning. In Proceedings of the IEEE/CVF international conference on computer vision, pages 7464–7473, 2019. 1
  48. 48.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. Sequence to sequence learning with neural networks. Advances in neural information processing systems, 27, 2014. 1, 2
  49. 49.Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216, 2019. 3
  50. 50.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 2
  51. 51.Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4566–4575, 2015. 6
  52. 52.Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko. Sequence to sequence-video to text. In Proceedings of the IEEE international conference on computer vision, pages 4534–4542, 2015. 2
  53. 53.Bairui Wang, Lin Ma, Wei Zhang, and Wei Liu. Reconstruction network for video captioning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7622–7631, 2018. 2
  54. 54.Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Luowei Zhou, Yucheng Zhao, Yujia Xie, Ce Liu, Yu-Gang Jiang, and Lu Yuan. Omnivl: One foundation model for image-language and video-language tasks. Advances in neural information processing systems, 35:5696–5710, 2022. 1
  55. 55.Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. Vatex: A large-scale, high-quality multilingual dataset for video-and-language research. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4581–4591, 2019. 3
  56. 56.Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al. Internvideo: General video foundation models via generative and discriminative learning. arXiv preprint arXiv:2212.03191, 2022. 7
  57. 57.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296, 2016. 3
  58. 58.Ran Xu, Caiming Xiong, Wei Chen, and Jason Corso. Jointly modeling deep video and compositional text to bridge vision and language in a unified framework. In Proceedings of the AAAI conference on artificial intelligence, 2015. 2
  59. 59.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Zero-shot video question answering via frozen bidirectional language models. ArXiv, abs/2206.08155, 2022. 7
  60. 60.Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid. Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10714–10726, 2023. 1, 2
  61. 61.Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Pal, Hugo Larochelle, and Aaron Courville. Describing videos by exploiting temporal structure. In Proceedings of the IEEE international conference on computer vision, pages 4507–4515, 2015. 2
  62. 62.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yi Zhou, Junyan Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qiang Qi, Ji Zhang, and Feiyan Huang. mplug-owl: Modularization empowers large language models with multimodality. ArXiv, abs/2304.14178, 2023. 7
  63. 63.Bowen Zhang, Hexiang Hu, and Fei Sha. Cross-modal and hierarchical modeling of video and text. In European Conference on Computer Vision, 2018. 3
  64. 64.Yue Zhao, Ishan Misra, Philipp Krahenb ü uhl, and Rohit Girdhar. Learning video representations from large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6586–6597, 2023. 2, 4, 5, 6, 7
  65. 65.Luowei Zhou, Chenliang Xu, and Jason Corso. Towards automatic learning of procedures from web instructional videos. In Proceedings of the AAAI Conference on Artificial Intelligence, 2018. 3
  66. 66.Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic. Cross-task weakly supervised learning from instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3537–3545, 2019. 3

Citation

MLA
Islam, M. M., et al. “Video ReCap: Recursive Captioning of Hour-Long Videos”. arXiv, 2024, http://arxiv.org/abs/2402.13250v6.
APA
Islam, M. M., Ho, N., Yang, X., Nagarajan, T., Torresani, L., & Bertasius, G. (2024). Video ReCap: Recursive Captioning of Hour-Long Videos. arXiv. http://arxiv.org/abs/2402.13250v6
Chicago
Islam, M. M., N. Ho, X. Yang, T. Nagarajan, L. Torresani, and G. Bertasius. 2024. “Video ReCap: Recursive Captioning of Hour-Long Videos”. arXiv. http://arxiv.org/abs/2402.13250v6.
Harvard
Islam, M.M. et al. (2024) “Video ReCap: Recursive Captioning of Hour-Long Videos”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.13250v6.
Vancouver
1. Islam MM, Ho N, Yang X, Nagarajan T, Torresani L, Bertasius G (2024) Video ReCap: Recursive Captioning of Hour-Long Videos. arXiv

BibTeX

@article{islam2024video,
  title = {Video ReCap: Recursive Captioning of Hour-Long Videos},
  author = {Islam, Md Mohaiminul and Ho, Ngan and Yang, Xitong and Nagarajan, Tushar and Torresani, Lorenzo and Bertasius, Gedas},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.13250v6},
  eprint = {2402.13250}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/