LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Haoning WuDongxu LiBei ChenJunnan Li

article2024NeurIPS484 citations

Establishes LongVideoBench, a challenging benchmark of hour-long interleaved video-subtitle inputs and a novel referring reasoning task that exposes substantial performance gaps in leading large multimodal models.

Listen

Artificial intelligence models are rapidly expanding their context windows to ingest large volumes of information, yet standard benchmarks primarily evaluate text-only inputs. Existing video benchmarks suffer from single-frame bias, where models achieve top scores by processing just a handful of frames or global summaries rather than analyzing extensive visual content over time. Consequently, practitioners lack reliable methods to measure whether vision-language models can genuinely process, retrieve, and reason across extended multimodal inputs such as hour-long subtitled videos.

The article introduces and evaluates LongVideoBench, a benchmark designed to assess the long-context multimodal understanding of large multimodal models. Its primary objective is to evaluate how effectively proprietary and open-source models retrieve fine-grained details and reason across extended, interleaved video and subtitle sequences ranging from seconds up to one hour in length.

To construct this evaluation, the authors compiled 3,763 diverse web videos across 10 categories, pairing them with aligned subtitles across four duration brackets up to 60 minutes. They developed a novel question-answering paradigm termed referring reasoning, creating 6,678 human-annotated multiple-choice questions divided into perception tasks (grounded in single video moments) and relation tasks (requiring temporal or sequential reasoning across multiple moments). The benchmark was then used to evaluate 22 models—comprising proprietary leaders, open-source long-context systems, image-based models, and video-specific architectures—under zero-shot conditions.

The evaluation revealed several key findings regarding model capabilities. First, unlike prior benchmarks, performance on LongVideoBench strictly improves as models process more frames; leading proprietary models such as GPT-4o and Gemini-1.5-Pro improved by over 10 percentage points on videos exceeding three minutes when scaled from 16 to 256 frames. Second, a substantial capability gap exists between proprietary and open-source systems: top proprietary models achieved accuracy between 60% and 67%, while open-source models scored around 40% to 53% and suffered performance degradations when pushed beyond 16 frames. Third, reasoning across temporal relationships proved significantly harder than single-scene perception, with sequence ordering tasks showing the lowest overall scores. Finally, model accuracy was unevenly distributed across video duration, dropping when target queries were located near the beginning or middle of videos rather than near the end.

These results demonstrate that existing open-source and specialized video models are not yet equipped for real-world long-context video comprehension, largely because prior training datasets over-indexed on short clips and high-level summaries. For decision-makers and developers, relying on current open-source systems for complex video analysis introduces operational risk and high error rates, whereas proprietary services offer viable but imperfect accuracy at higher computational and API costs.

Organizations developing or deploying multimodal systems should prioritize training regimens that emphasize fine-grained multimodal retrieval and temporal reasoning rather than simple global captioning. For production deployments requiring hour-long video analysis, practitioners should deploy high-capacity models capable of processing dense frame inputs while monitoring for retrieval failures at earlier timestamps. While the benchmark provides high confidence through human-verified annotations and consistent test-validation results, current findings are limited to vision and text inputs under one hour and do not yet evaluate the audio modality.

Cover for LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Abstract

Large multimodal models (LMMs) are processing increasingly longer and richer inputs. Albeit the progress, few public benchmark is available to measure such development. To mitigate this gap, we introduce LongVideoBench, a question-answering benchmark that features video-language interleaved inputs up to an hour long. Our benchmark includes 3,763 varying-length web-collected videos with their subtitles across diverse themes, designed to comprehensively evaluate LMMs on long-term multimodal understanding. To achieve this, we interpret the primary challenge as to accurately retrieve and reason over detailed multimodal information from long inputs. As such, we formulate a novel video question-answering task termed referring reasoning. Specifically, as part of the question, it contains a referring query that references related video contexts, called referred context. The model is then required to reason over relevant video details from the referred context. Following the paradigm of referring reasoning, we curate 6,678 human-annotated multiple-choice questions in 17 fine-grained categories, establishing one of the most comprehensive benchmarks for long-form video understanding. Evaluations suggest that the LongVideoBench presents significant challenges even for the most advanced proprietary models (e.g. GPT-4o, Gemini-1.5-Pro, GPT-4-Turbo), while their open-source counterparts show an even larger performance gap. In addition, our results indicate that model performance on the benchmark improves only when they are capable of processing more frames, positioning LongVideoBench as a valuable benchmark for evaluating future-generation long-context LMMs.

Table of Contents

  • 1 Introduction
  • 2 The Referring Reasoning Task
  • 3 Dataset Construction
  • 3.1 Groups of Videos
  • 3.2 Video and Subtitle Collection
  • 3.3 Annotating Questions and Answers
  • 4 Evaluation of LongVideoBench
  • 4.1 Models and Evaluation Strategies
  • 4.2 Main Results
  • 4.3 Leaderboard
  • 4.4 Performance w.r.t. Referring Query Depth
  • 5 Related Works
  • 6 Conclusion
  • References
  • A Additional Experimental Settings
  • A.1 A Brief Introduction on Participating Models
  • A.2 Prompts and Settings
  • B More Visualizations w.r.t. Referring Query Depth
  • C Limitations
  • D Broader Impacts
  • E Details about Human Annotation
  • E.1 Instructions and Annotation Interface
  • E.2 Hourly Wage and Total Compensation
  • E.3 IRB Approval for Human Study
  • F URLs to Websites
  • G Author Statement on Responsibility
  • H Hosting, Licensing and Maintenance Plan
  • I Dataset Sheet for LongVideoBench

Knowls

  1. Knowl 1 — Referring Reasoning Task Formulation

    definition

    The referring reasoning task is a video question-answering formulation designed to evaluate long-context multimodal understanding in Large Multimodal Models (LMMs) while overcoming single-frame bias. Instead of posing summary or global overview questions that can be solved using a few isolated key frames, referring reasoning requires models to retrieve granular multimodal evidence and reason over contextual interconnections across long video sequences.

    A referring reasoning instance consists of three core components:

    1. Referring Query: A text segment that pinpoints one or more specific moments in the video (composed of video frames and/or subtitle transcripts), known as the referred context.
    2. Question Body: A query that asks about specific visual, semantic, or relational details within the referred context.
    3. Multiple-Choice Candidates: A set of options comprising one correct answer and 3 to 4 plausible distracting options.

    The task is organized into two operational tiers:

    • Level 1 (Perception): The referring query identifies a single video moment, and the question evaluates visual perception of a specific concept (e.g., object identification, attribute recognition, action/event observation) occurring in that moment.
    • Level 2 (Relation): The referred context spans multiple video moments (either linked by temporal order—before, after, concurrent—or sharing a common entity). The question evaluates relational reasoning across these moments, such as tracking entities, detecting attribute changes across time, or determining temporal event sequences.
  2. Knowl 2 — LongVideoBench Dataset Composition and Duration Distribution

    definition

    LongVideoBench is a video question-answering benchmark comprising 3,763 web-collected videos with aligned subtitles and 6,678 human-annotated multiple-choice questions (MCQs). Across the dataset, the average video duration is 473 seconds, the average question length is 43.53 words, and the average answer length is 8.28 words.

    Videos are drawn from 119 channels across 10 diverse themes:

    • Life-related (3): Travel Guides (LT), Life Vlogs (LV), Cooking/Recipes (LC)
    • Knowledge-related (5): Art (KA), History (KH), Geography (KG), STEM (KS), Computer Science (KC)
    • Entertainment & Media (2): Movie Recaps (MR), News Programs (NP)

    The dataset is partitioned into four progressive duration groups:

    • (8s,15s](8\text{s}, 15\text{s}]: 884 videos (546 landscape, avg. duration 11.06s; 338 portrait, avg. duration 11.93s)
    • (15s,60s](15\text{s}, 60\text{s}]: 925 videos (551 landscape, avg. duration 33.88s; 374 portrait, avg. duration 38.59s)
    • (180s,600s](180\text{s}, 600\text{s}]: 986 landscape videos (avg. duration 389s)
    • (900s,3600s](900\text{s}, 3600\text{s}]: 966 landscape videos (avg. duration 1,408s)

    The benchmark is divided into two distinct splits:

    • Validation Set: 752 videos and 1,337 MCQs (20% of total data) with public answer labels.
    • Test Set: 3,011 videos and 5,341 MCQs (80% of total data) with hidden ground-truth labels to prevent benchmark contamination.
  3. Knowl 3 — Taxonomy of the 17 Fine-Grained Referring Reasoning Categories

    definition

    LongVideoBench categorizes its 6,678 questions into 17 fine-grained tasks across two complexity levels based on the referring query modality and the target answer type:

    Perception Level (L1, 3,204 Questions):

    1. Scene-Referred Event (S2E, 410 Qs): Query specifies a visual scene; answer targets an event occurring in that scene.
    2. Scene-Referred Object Existence (S2O, 403 Qs): Query specifies a visual scene; answer targets an object existing in that scene.
    3. Scene-Referred Object Attribute (S2A, 403 Qs): Query specifies a visual scene q1q_1 and object q2q_2; answer targets a visual attribute of q2q_2 in q1q_1.
    4. Event-Referred Object (E2O, 393 Qs): Query specifies an event; answer targets an object/person participating in that event.
    5. Object-Referred Event (O2E, 401 Qs): Query specifies an object/person; answer targets an event occurring during their appearance.
    6. Text-Referred Event (T2E, 398 Qs): Query specifies a subtitle segment; answer targets an event occurring concurrently with that subtitle.
    7. Text-Referred Object Existence (T2O, 387 Qs): Query specifies a subtitle segment; answer targets an object present while that subtitle appears.
    8. Text-Referred Object Attribute (T2A, 402 Qs): Query specifies subtitle q1q_1 and object q2q_2; answer targets an attribute of q2q_2 while q1q_1 appears.

    Relation Level (L2, 3,474 Questions): 9. Event Before/After Event (E3E, 406 Qs): Query specifies an event; answer targets an event preceding or succeeding it. 10. Object Before/After Object (O3O, 394 Qs): Query specifies an entity; answer targets an entity appearing before or after it. 11. Sequence of Scenes (SSS, 398 Qs): Query specifies multiple scenes (>3>3); answer targets the correct chronological sequence order among the scenes (distractors are permutations of the sequence). 12. Scene-Referred Object Tracking (SOS, 381 Qs): Query specifies scene q1q_1 and object q2q_2; answer targets another scene in which q2q_2 appears. 13. Scene-Referred Object Attribute Change (SAA, 375 Qs): Query specifies two scenes q1,q2q_1, q_2 and object q3q_3; answer targets the attribute change of q3q_3 from q1q_1 to q2q_2. 14. Event Before/After Text (T3E, 401 Qs): Query specifies a subtitle; answer targets an event occurring before or after that subtitle. 15. Object Before/After Text (T3O, 391 Qs): Query specifies a subtitle; answer targets an entity appearing before or after that subtitle. 16. Text-Referred Object Tracking (TOS, 380 Qs): Query specifies scene q1q_1 and object q2q_2; answer targets the specific subtitle spoken during an appearance of q2q_2 other than in q1q_1. 17. Text-Referred Object Attribute Change (TAA, 348 Qs): Query specifies two subtitles q1,q2q_1, q_2 and object q3q_3; answer targets the attribute change of q3q_3 from q1q_1 to q2q_2.

  4. Knowl 4 — Interleaved Video-Language Input Representation

    model/method

    To evaluate multimodal models in a manner that mirrors human viewing of subtitled video, inputs in LongVideoBench are formatted as temporally-aligned interleaved multimodal sequences combining video frames and subtitle transcriptions.

    For a given video with sampled frame timestamps {tf}\{t_f\} and transcribed/original subtitle chunks {Sk}\{S_k\} with midpoint timestamps tmid(Sk)t_{\text{mid}}(S_k):

    1. Each subtitle chunk SkS_k is inserted directly in-between the two consecutive video frames fif_i and fi+1f_{i+1} whose timestamps satisfy tfi≤tmid(Sk)<tfi+1t_{f_i} \le t_{\text{mid}}(S_k) < t_{f_{i+1}}.
    2. The interleaved sequence is fed into the Large Multimodal Model as an ordered list of image placeholders (or image embedding tokens) and text tokens.
    3. The interleaved video-subtitle context is followed by the question stem, multiple candidate options labeled with letters (e.g., A, B, C, D, E), and the fixed suffix prompt:
    Question: <Question>
    <Options>
    Answer with the option's letter from the given choices directly.
    
  5. Knowl 5 — Scaling Behavior of LMMs with Input Frame Budget on Long Videos

    data/table

    Evaluation of long-context LMMs on the LongVideoBench validation set (1,337 MCQs) across varying maximum input frame allocations (1,4,8,16,32,64,128,2561, 4, 8, 16, 32, 64, 128, 256 frames capped at 1 fps) reveals a fundamental divide between proprietary and open-source models. Proprietary models consistently scale with higher frame counts, achieving their highest accuracy on long videos at 256 frames, whereas open-source models saturate or degrade beyond 16 frames.

    Model 1 Frame 4 Frames 8 Frames 16 Frames 32 Frames 64 Frames 128 Frames 256 Frames
    GPT-4o (0513) 41.7 48.7 53.3 58.0 58.5 62.0 63.5 66.7
    Gemini-1.5-Pro (0514) 38.6 44.9 51.0 52.7 55.2 58.6 61.9 64.0
    Gemini-1.5-Flash (0514) 38.1 45.2 48.9 50.8 53.5 56.8 58.9 61.6
    GPT-4-Turbo (0409) 43.2 48.1 49.7 52.7 54.0 56.0 57.5 59.0
    Idefics2 41.5 46.4 48.5 49.7 47.8 30.9 – –
    Phi-3-Vision-Instruct 40.8 47.5 48.1 49.6 49.1 Context Exceeded – –
    Mantis-Idefics2 38.7 43.5 46.1 47.0 45.4 30.2 – –
    Mantis-BakLLaVA 38.7 42.0 43.0 43.7 42.8 Context Exceeded – –

    Key observations:

    • On videos longer than 180 seconds, GPT-4o and Gemini-1.5-Pro improve by over 10%10\% when increasing input frames from 16 to 256 (e.g., GPT-4o gains from 53.8%53.8\% to 69.1%69.1\% on (180s,600s](180\text{s}, 600\text{s}] and from 52.2%52.2\% to 60.9%60.9\% on (900s,3600s](900\text{s}, 3600\text{s}]).
    • Open-source models (Idefics2, Mantis-Idefics2) experience severe performance degradation at 64 frames (dropping to ≈30.2%−30.9%\approx 30.2\% - 30.9\%) before reaching context limits, demonstrating deficiencies in long-context multimodal processing.
  6. Knowl 6 — LongVideoBench Test Leaderboard Across 23 Large Multimodal Models

    data/table

    Evaluation of 23 Large Multimodal Models on the LongVideoBench test set (5,341 MCQs) shows that proprietary long-context LMMs substantially outperform open-source counterparts. Scaling the language model backbone yields substantial improvements for open-source architectures.

    Model Val Total Test Total (8, 15]s (15, 60]s (180, 600]s (900, 3600]s
    Proprietary Long-Context LMMs
    GPT-4o (0513) 66.7 66.7 71.6 76.8 66.7 61.6
    Gemini-1.5-Pro (0514) 64.0 64.4 70.2 75.3 65.0 59.1
    Gemini-1.5-Flash (0514) 61.6 62.4 66.1 73.1 63.1 57.3
    GPT-4-Turbo (0409) 59.1 60.7 66.4 71.1 61.7 54.5
    Open-Source Long-Context LMMs
    Phi-3-Vision-Instruct 49.6 49.9 58.3 59.6 48.4 45.1
    Idefics2 49.7 49.4 57.4 60.4 47.3 44.7
    Mantis-Idefics2 47.0 47.6 56.1 61.4 44.6 42.5
    Mantis-BakLLaVA 43.7 43.7 51.3 52.7 41.1 40.1
    Open-Source Video LMMs
    PLLaVA-34B 53.2 53.5 60.1 66.8 50.8 49.1
    LLaVA-Next-Video-34B 50.5 50.5 57.6 61.6 48.7 45.9
    PLLaVA-13B 45.6 45.1 52.9 54.3 42.9 41.2
    LLaVA-Next-Video-M7B 43.5 43.5 50.9 53.1 42.6 38.9
    ShareGPT4Video 39.7 41.8 46.9 50.1 40.0 38.7
    PLLaVA-7B 40.2 39.2 45.3 47.3 38.5 35.2
    VideoChat2 (Mistral-7B) 39.3 41.2 49.3 49.3 39.0 37.5
    VideoLLaVA 39.1 37.6 43.1 44.6 36.4 34.4
    VideoChat2 (Vicuna-7B) 36.0 35.1 38.1 40.5 33.5 33.6
    Open-Source Multi-Image LMMs (8 frames)
    LLaVA-Next-Mistral-7B 49.1 47.1 53.4 57.2 46.9 42.1
    InstructBLIP-T5-XXL 43.3 43.8 48.1 50.1 44.5 40.0
    BLIP-2-T5-XXL 42.7 43.5 46.7 47.4 44.2 40.9
    LLaVA-1.5-13B 43.4 43.1 49.0 51.1 41.8 39.6
    LLaVA-1.5-7B 40.3 40.4 45.0 47.4 40.1 37.0
    mPLUG-Owl2 39.1 39.4 49.4 47.3 38.7 34.3

    Notable findings include:

    1. Open-source video LMMs show no clear advantage over multi-image models; under the same architecture, LLaVA-Next-Mistral-7B (47.1%47.1\%) outperforms video-tuned LLaVA-Next-Video-M7B (43.5%43.5\%).
    2. LLM backbone scaling provides substantial improvements: PLLaVA-34B (53.5%53.5\%) improves upon PLLaVA-13B (45.1%45.1\%) by +8.4%+8.4\% and PLLaVA-7B (39.2%39.2\%) by +14.3%+14.3\%.
  7. Knowl 7 — Multimodal Input Ablation: Impact of Video Frames vs. Subtitles

    empirical result

    An ablation study on the validation set of LongVideoBench evaluates the relative contribution of video frames and text subtitles across 8 long-context LMMs:

    Frames? Subtitles? GPT-4o Gemini-Pro GPT-4-T Gemini-Flash Idefics2 Phi-3-V M-Idefics2 M-BakLLaVA
    ✓ 44.6 43.0 45.2 39.2 25.6 40.7 31.7 31.1
    ✓ 60.6 62.9 56.0 60.2 49.4 49.5 45.8 43.5
    ✓ ✓ 66.7 63.9 59.0 61.6 49.7 49.6 47.0 43.7

    Key takeaways:

    1. Visual Modality is Essential: Removing video frames and using subtitles alone causes sharp accuracy drops across all models (e.g., GPT-4o falls from 66.7%66.7\% to 44.6%44.6\%, Gemini-1.5-Pro from 63.9%63.9\% to 43.0%43.0\%, and Idefics2 from 49.7%49.7\% to 25.6%25.6\%).
    2. Proprietary vs. Open-Source Modality Integration: Proprietary models effectively synthesize interleaved visual and textual modalities, gaining +1.0%+1.0\% to +6.1%+6.1\% accuracy over visual-only inputs. In contrast, open-source models show near-zero gain when adding subtitles to frames (+0.1%+0.1\% to +1.2%+1.2\%), demonstrating a failure to effectively integrate interleaved multimodal information.
  8. Knowl 8 — Performance Disparity Between Perception (L1) and Relation (L2) Tasks

    empirical result

    Across all evaluated models in LongVideoBench, there is a consistent performance gap between Level 1 (Perception) and Level 2 (Relation) questions, demonstrating that multi-moment relational and temporal reasoning is significantly more difficult than single-moment visual grounding.

    On the test set:

    • GPT-4o: Scores 71.6%71.6\% on L1 Perception (e.g., S2E: 76.8%76.8\%, T2A: 77.2%77.2\%) versus 62.6%62.6\% on L2 Relation (e.g., SSS: 44.3%44.3\%, SOS: 75.6%75.6\%).
    • Gemini-1.5-Pro: Scores 70.2%70.2\% on L1 Perception versus 59.0%59.0\% on L2 Relation.
    • PLLaVA-34B (top open-source): Scores 60.1%60.1\% on L1 Perception versus 47.9%47.9\% on L2 Relation.

    The single hardest category across the benchmark is Sequence of Scenes (SSS), where the model must determine the correct temporal ordering of three or more scenes from candidate permutations. All models perform worst on SSS:

    • GPT-4o: 44.3%44.3\%
    • Gemini-1.5-Flash: 43.1%43.1\%
    • GPT-4-Turbo: 44.8%44.8\%
    • Gemini-1.5-Pro: 55.2%55.2\%
    • Open-source models range between 10.6%10.6\% (VideoChat2 Mistral-7B) and 44.2%44.2\% (PLLaVA-34B).

    This discrepancy demonstrates that current LMMs lack robust temporal comprehension and struggle to maintain precise chronological ordering across long video contexts.

  9. Knowl 9 — Impact of Referring Query Temporal Depth on Model Accuracy

    empirical result

    Analyzing question accuracy as a function of the temporal location (depth) of the referred context within a video demonstrates systematic positional performance degradation in LMMs, reflecting the multimodal counterpart of the 'needle in a haystack' (NIAH) phenomenon:

    1. Longer Token-Distance Degradation: All evaluated models perform worse when the queried moment is located near the beginning of the video (i.e., at a greater relative token distance from the question prompt presented at the end of the input sequence).
    2. Lost-in-the-Middle Effect: Queries targeting evidence located in the middle of long videos present greater challenge to LMMs than those targeting boundaries.
    3. Duration Amplification: The performance drop associated with query depth becomes increasingly severe as the total video duration expands from short clips ((8s,60s](8\text{s}, 60\text{s}]) to hour-long inputs ((900s,3600s](900\text{s}, 3600\text{s}]).
  10. Knowl 10 — Limitations of LongVideoBench

    limitation

    LongVideoBench has two explicitly stated scope limitations:

    1. Modality Constraints: The benchmark evaluates only vision (video frames) and text (original or transcribed subtitles) modalities, omitting raw audio data (such as non-verbal acoustic events, background audio, tone of voice, and ambient sound).
    2. Video Duration Cap: The maximum video duration in the benchmark is restricted to 1 hour (3,600 seconds), leaving ultra-long video understanding beyond 1 hour unmeasured.

Coverage note — No substantial contributed material was omitted from the extraction.

References

  1. 1.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models, 2023.
  2. 2.OpenAI. Hello gpt-4o. https://openai.com/index/hello-gpt-4o/, May 2024a.
  3. 3.Gemini Team. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024.
  4. 4.Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. Ruler: What’s the real context size of your long-context language models?, 2024.
  5. 5.Chonghua Wang, Haodong Duan, Songyang Zhang, Dahua Lin, and Kai Chen. Ada-leval: Evaluating long-context llms with length-adaptable benchmarks, 2024a.
  6. 6.gkamradt. Llmtest_needleinahaystack, 2024. URL https://github.com/gkamradt/LLMTest_NeedleInAHaystack.
  7. 7.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video question answering via gradually refined attention over appearance and motion. In ACM Multimedia, 2017.
  8. 8.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. Activitynet-qa: A dataset for understanding complex web videos via question answering. arXiv preprint arXiv:1906.02467, 2019. URL https://arxiv.org/abs/1906.02467.
  9. 9.Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. Next-qa:next phase of question-answering to explaining temporal actions, 2021.
  10. 10.Gaoang Wang, Jenq-Neng Hwang, Yanting Zhang, and Yan Lu. Mv-bench: A benchmark for multi-view video understanding. arXiv preprint arXiv:2308.12345, 2023. URL https://arxiv.org/abs/arXiv.2308.12345.
  11. 11.Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A diagnostic benchmark for very long-form video language understanding. Advances in Neural Information Processing Systems, 2023. URL https://proceedings.neurips.cc/paper_files/paper/2023/hash/90ce332aff156b910b002ce4e6880dec-Abstract-Datasets_and_Benchmarks.html. arXiv preprint arXiv:2308.09126.
  12. 12.Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, Yan Lu, Jenq-Neng Hwang, and Gaoang Wang. Moviechat: From dense token to sparse memory for long video understanding. arXiv preprint arXiv:2307.16449, 2023. URL https://arxiv.org/abs/2307.16449.
  13. 13.Hongjie Zhang, Yi Liu, Lu Dong, Yifei Huang, Zhenhua Ling, Yali Wang, Limin Wang, and Yu Qiao. Movqa: A benchmark of versatile question-answering for long-form movie understanding. arXiv preprint arXiv:2312.04817, 2023a. URL https://arxiv.org/pdf/2312.04817.
  14. 14.OpenAI. Whisper-v3-large: A large scale multimodal model. https://github.com/openai/whisper-v3-large, 2024b.
  15. 15.Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, Qiong Yan, Xiongkuo Min, Guangtao Zhai, and Weisi Lin. Q-align: Teaching lmms for visual scoring via discrete text-defined levels. In ICML2024, 2024.
  16. 16.Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Jilan Xu, Zun Wang, Yansong Shi, Tianxiang Jiang, Songze Li, Hongjie Zhang, Yifei Huang, Yu Qiao, Yali Wang, and Limin Wang. Internvideo2: Scaling video foundation models for multimodal video understanding, 2024b.
  17. 17.Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding, 2023b.
  18. 18.Qiang Zhang, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424, 2023c. URL https://arxiv.org/abs/2306.05424.
  19. 19.Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng. Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024.
  20. 20.KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding, 2023a.
  21. 21.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023b.
  22. 22.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. CoRR, abs/2310.03744, 2023a.
  23. 23.Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Haowei Liu, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. mPLUG-Owl2: Revolutionizing multi-modal large language model with modality collaboration. CoRR, abs/2311.04257, 2023.
  24. 24.Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models, 2023c.
  25. 25.Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, and Ser-Nam Lim. Ma-lmm: Memory-augmented large multimodal model for long-term video understanding, 2024.
  26. 26.Reuben Tan, Ximeng Sun, Ping Hu, Jui hsien Wang, Hanieh Deilamsalehy, Bryan A. Plummer, Bryan Russell, and Kate Saenko. Koala: Key frame-conditioned long video-llm, 2024.
  27. 27.Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel. World model on million-length video and language with blockwise ringattention, 2024a.
  28. 28.Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Caio César Teodoro Mendes, Weizhu Chen, Vishrav Chaudhary, Parul Chopra, Allie Del Giorno, Gustavo de Rosa, Matthew Dixon, Ronen Eldan, Dan Iter, Amit Garg, Abhishek Goswami, Suriya Gunasekar, Emman Haider, Junheng Hao, Russell J. Hewett, Jamie Huynh, Mojan Javaheripi, Xin Jin, Piero Kauffmann, Nikos Karampatziakis, Dongwoo Kim, Mahoud Khademi, Lev Kurilenko, James R. Lee, Yin Tat Lee, Yuanzhi Li, Chen Liang, Weishung Liu, Eric Lin, Zeqi Lin, Piyush Madan, Arindam Mitra, Hardik Modi, Anh Nguyen, Brandon Norick, Barun Patra, Daniel Perez-Becker, Thomas Portet, Reid Pryzant, Heyang Qin, Marko Radmilac, Corby Rosset, Sambudha Roy, Olatunji Ruwase, Olli Saarikivi, Amin Saied, Adil Salim, Michael Santacroce, Shital Shah, Ning Shang, Hiteshi Sharma, Xia Song, Masahiro Tanaka, Xin Wang, Rachel Ward, Guanhua Wang, Philipp Witte, Michael Wyatt, Can Xu, Jiahang Xu, Sonali Yadav, Fan Yang, Ziyi Yang, Donghan Yu, Chengruidong Zhang, Cyril Zhang, Jianwen Zhang, Li Lyna Zhang, Yi Zhang, Yue Zhang, Yunan Zhang, and Xiren Zhou. Phi-3 technical report: A highly capable language model locally on your phone, 2024.
  29. 29.Hugo Laurençon, Léo Tronchon, Matthieu Cord, and Victor Sanh. What matters when building vision-language models?, 2024.
  30. 30.Huggingface. Introducing idefics: An open reproduction of state-of-the-art visual language model, 2023. URL https://huggingface.co/blog/idefics.
  31. 31.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7b, 2023.
  32. 32.Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, and Wenhu Chen. Mantis: Interleaved multi-image instruction tuning, 2024.
  33. 33.SkunkworksAI. BakLLaVA, 2024. URL https://github.com/SkunkworksAI/BakLLaVA.
  34. 34.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. CoRR, abs/2305.06500, 2023.
  35. 35.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. CoRR, abs/2304.08485, 2023b.
  36. 36.Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Lava-next: Improved reasoning, ocr, and world knowledge, 2024b. URL https://llava-vl.github.io/blog/2024-01-30-llava-next/.
  37. 37.Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual representation by alignment before projection, 2023.
  38. 38.Lin Chen, Jisong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. Sharegpt4v: Improving large multi-modal models with better captions. arXiv preprint arXiv:2311.12793, 2023.
  39. 39.Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, Feng Zhao, and Jiaqi Wang. Sharegpt4video: Improving video understanding and generation with better captions, 2024. URL https://arxiv.org/abs/2406.04325.
  40. 40.Yuanhan Zhang, Bo Li, Haotian Liu, Yong Jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li. Llava-next: A strong zero-shot video understanding model, 2024. URL https://llava-vl.github.io/blog/2024-04-30-llava-next-video/.

Citation

MLA
Wu, H., et al. “LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding”. arXiv, 2024, http://arxiv.org/abs/2407.15754v1.
APA
Wu, H., Li, D., Chen, B., & Li, J. (2024). LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding. arXiv. http://arxiv.org/abs/2407.15754v1
Chicago
Wu, H., D. Li, B. Chen, and J. Li. 2024. “LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding”. arXiv. http://arxiv.org/abs/2407.15754v1.
Harvard
Wu, H. et al. (2024) “LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2407.15754v1.
Vancouver
1. Wu H, Li D, Chen B, Li J (2024) LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding. arXiv

BibTeX

@article{wu2024longvideobench,
  title = {LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding},
  author = {Wu, Haoning and Li, Dongxu and Chen, Bei and Li, Junnan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2407.15754v1},
  eprint = {2407.15754}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors