MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Kunchang LiYali WangYinan HeYizhuo LiYi WangYi LiuZun WangJilan XuGuo ChenPing Luo

article2023CVPR1,346 citations

Presents MVBench, a comprehensive benchmark comprising 20 dynamic video tasks designed to evaluate temporal reasoning in multimodal language models beyond static image perception, alongside VideoChat2, a progressive baseline that outperforms prior methods by over 15%.

Listen

Recent advances in artificial intelligence have produced multi-modal large language models capable of processing visual and textual data together. However, existing evaluation benchmarks primarily test static image comprehension, neglecting how well models understand changes over time in dynamic video. Current video benchmarks remain narrow, expensive to build, or vulnerable to subjective automated scoring. This creates a critical assessment gap for organizations seeking to deploy artificial intelligence systems in dynamic, real-world environments.

The article addresses this gap by introducing MVBench, a comprehensive benchmark designed to evaluate temporal video understanding across diverse tasks, and VideoChat2, an open-source baseline model engineered specifically for video comprehension.

The researchers developed MVBench by transforming static image understanding tasks into 20 distinct dynamic video tasks spanning perception and cognition, such as action sequence tracking, object trajectory, and counterfactual inference. They established an automated pipeline using eleven public video datasets to generate multiple-choice question-and-answer pairs, ensuring standardized evaluation without costly manual labeling or biased scoring. To advance model performance, the authors also developed VideoChat2 using a 2-million-sample instruction dataset across 34 sources and a three-stage progressive training strategy that aligns visual features with language representations.

Evaluation results show that existing multi-modal models perform poorly on dynamic video tasks. Many models scored near random guessing, with leading systems achieving around 35.5% accuracy—barely outperforming text-only baselines without video input. In contrast, the standard VideoChat2 model achieved 51.1% accuracy, outperforming existing models by more than 15 percentage points. An upgraded version based on the Mistral architecture reached 60.4% accuracy, exceeding commercial proprietary models such as GPT-4V (43.5%) by 16.9 percentage points. Ablation analyses revealed that performance gains are heavily driven by the quality of the visual encoder and instruction-tuning data rather than language model size alone.

These findings demonstrate that static image perception does not translate to dynamic temporal reasoning. For technical leaders and decision-makers, relying on standard vision-language models for video-based operational tasks presents substantial reliability risks. Building effective video artificial intelligence requires specialized temporal architectures and diverse video training rather than merely scaling traditional language models.

Organizations should adopt comprehensive temporal benchmarks like MVBench before deploying computer vision models in operational environments requiring sequence or event reasoning. Model developers should prioritize dedicated video pre-training backbones and diverse instruction tuning. Future development must focus on improving fine-grained localization, counting, and integrating additional modalities such as audio and subtitles, where current models continue to face limitations.

Cover for MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Abstract

With the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models. However, most benchmarks predominantly assess spatial understanding in the static image tasks, while overlooking temporal understanding in the dynamic video tasks. To alleviate this issue, we introduce a comprehensive Multi-modal Video understanding Benchmark, namely MVBench, which covers 20 challenging video tasks that cannot be effectively solved with a single frame. Specifically, we first introduce a novel static-to-dynamic method to define these temporal-related tasks. By transforming various static tasks into dynamic ones, we enable the systematic generation of video tasks that require a broad spectrum of temporal skills, ranging from perception to cognition. Then, guided by the task definition, we automatically convert public video annotations into multiple-choice QA to evaluate each task. On one hand, such a distinct paradigm allows us to build MVBench efficiently, without much manual intervention. On the other hand, it guarantees evaluation fairness with ground-truth video annotations, avoiding the biased scoring of LLMs. Moreover, we further develop a robust video MLLM baseline, i.e., VideoChat2, by progressive multi-modal training with diverse instruction-tuning data. The extensive results on our MVBench reveal that, the existing MLLMs are far from satisfactory in temporal understanding, while our VideoChat2 largely surpasses these leading models by over 15% on MVBench. All models and data are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 MVBench
  • 3.1 Temporal Task Definition
  • 3.2 Automatic QA Generation
  • 3.3 Prompt Design for Evaluation
  • 4 VideoChat2
  • 4.1 Instruction-Tuning Data
  • 4.2 Progressive Multi-Modal Training
  • 5 Experiments
  • 5.1 Results on MVBench
  • 5.2 More Comparisons
  • 5.3 Ablations of VideoChat2
  • 6 Conclusion
  • A Training Hyperparameters
  • B More Ablations
  • C Details of QA Generation
  • D Results on Challenging Video QA
  • E Leaderboards and Analyses
  • F Qualitative Results
  • References

Knowls

  1. Knowl 1 — MVBench Temporal Task Definition and Taxonomy

    definition

    The Multi-modal Video understanding Benchmark (MVBench) systematically defines 20 temporal video understanding tasks by adapting 9 traditional static image understanding categories into dynamic video tasks that require multi-frame temporal reasoning:

    1. Action Tasks:

      • Action Sequence: Identifying events occurring immediately before or after a target action.
      • Action Prediction: Predicting subsequent actions based on initial video context.
      • Action Antonym: Differentiating a correct action from its reverse/inversely ordered counterpart.
      • Fine-grained Action: Distinguishing fine-grained actions among visually similar categories.
      • Unexpected Action: Recognizing surprising, humorous, or unexpected actions.
    2. Object Tasks:

      • Object Existence: Determining if a specific object exists during a specific temporal event.
      • Object Interaction: Identifying objects engaged in a particular action.
      • Object Shuffle: Tracking the final position of an object undergoing occluded shuffles.
    3. Position Tasks:

      • Moving Direction: Identifying the movement trajectory (e.g., up-right, down-left, stationary) of an object.
      • Action Localization: Determining the temporal segment (start, middle, end, throughout) during which an action occurs.
    4. Scene Tasks:

      • Scene Transition: Identifying source and destination environments during scene changes.
    5. Count Tasks:

      • Action Count: Counting occurrences of a specific repeated action.
      • Moving Count: Counting the number of distinct objects performing a specific motion.
    6. Attribute Tasks:

      • Moving Attribute: Recognizing visual properties (color, shape, material) of an object while in motion.
      • State Change: Detecting physical state alterations of objects over time.
    7. Pose Tasks:

      • Fine-grained Pose: Identifying subtle human posture changes over time.
    8. Character Tasks:

      • Character Order: Recognizing the chronological sequence in which written letters appear.
    9. Cognition Tasks:

      • Egocentric Navigation: Forecasting subsequent navigation steps from first-person viewpoints.
      • Episodic Reasoning: Inferring high-level narrative motivations and intent in multi-character video episodes.
      • Counterfactual Inference: Reasoning about hypothetical alternative outcomes if specific collision/interaction events were modified.
  2. Knowl 2 — MVBench Automated Data Curation and QA Generation Pipeline

    model/method

    MVBench constructs a 20-task benchmark containing 4,000 multiple-choice Question-Answering (QA) pairs (200 pairs per task) across 11 source video datasets: STAR, PAXION, Moments in Time V1, FunQA, CLEVRER, Perception Test, Charades-STA, MoVQA, NTU RGB+D, VLN-CE, and TVQA.

    The curation pipeline consists of three processing steps:

    1. Data Filtration:

      • Video Diversity: Collects videos across indoor/outdoor scenes and first-person/third-person viewpoints.
      • Temporal Sensitivity: Retains clips of intermediate duration (primarily 5 s5\,\text{s} to 35 s35\,\text{s}) to avoid trivial static clips and overly complex long contexts.
      • Difficulty Balancing: Shifts start/end boundaries in STAR to increase localization complexity, and filters CLEVRER questions requiring more than 10 object-property conditions.
    2. QA Generation:

      • Template-Based Construction: Constructs option candidates directly from structured annotations (e.g., pairing ground-truth actions with reversed actions and "not sure" in Action Antonym; generating 4 directional vectors plus "stationary" in Moving Direction).
      • LLM-Based Generation: Prompts ChatGPT using formal task definitions to generate 3 to 5 multiple-choice questions from open-ended annotations (such as FunQA and MoVQA).
    3. Option Processing:

      • Samples 3 to 5 candidate options per question with randomized option order.
      • Enforces uniform option lengths using LLM-based paraphrasing to prevent models from exploiting text-length heuristics.
  3. Knowl 3 — Evaluation Prompt Design and Constrained Option Extraction for Video MLLMs

    model/method

    MVBench evaluates Multi-modal Large Language Models (MLLMs) using a two-part prompt strategy that eliminates external parsing models and achieves a 100%100\% option extraction rate:

    1. System Prompt: Enforces scrutiny of temporal dynamics:

      "Carefully watch the video and pay attention to the cause and sequence of events, the detail and movement of objects and the action and pose of persons. Based on your observations, select the best option that accurately addresses the question."

    2. Answer Formatting Prompt: Options are enclosed in parenthesized keys in the question, followed by the suffix prompt:

      "Best Option: ("

    Appending "Best Option: (" constrains the MLLM to output the choice identifier directly, eliminating parsing failures. On MVBench, this prompt improves extraction hit ratios from 78.2%78.2\% to 100%100\% for VideoChat (raising accuracy from 22.8%22.8\% to 35.5%35.5\%) and from 64.6%64.6\% to 100%100\% for VideoChatGPT (raising accuracy from 22.0%22.0\% to 32.8%32.8\%).

  4. Knowl 4 — VideoChat2 Progressive Multi-Modal Training Architecture

    model/method

    VideoChat2 is a video Multi-modal Large Language Model composed of an Unmasked Teacher (UMT-L) vision encoder, a BERT-base initialized Q-Former, a linear projection layer, and a pretrained Large Language Model (Vicuna-7B v0 or Mistral-7B).

    Training follows a three-stage progressive paradigm:

    1. Stage 1: Vision-Language Alignment:

      • UMT-L is frozen. Q-Former uses 32 learnable query tokens to compress visual features.
      • Trained on 15M image captions (CC3M, CC12M) and 10M video captions (WebVid-10M) using Vision-Text Contrastive (VTC), Vision-Text Matching (VTM), and Vision-grounded Text Generation (VTG) losses on 4-frame video inputs for 10 epochs.
    2. Stage 2: Vision-Language Connection:

      • UMT-L is unfrozen alongside Q-Former; 64 additional randomly initialized query tokens are added (totaling 96 queries).
      • Query representations are mapped via a linear projection into the LLM embedding space.
      • Trained with VTG loss on 2M image captions (COCO, Visual Genome, SBU) and 10M video captions (InternVid) on 4-frame inputs for 1 epoch.
    3. Stage 3: Multi-Modal Instruction Tuning:

      • The LLM is adapted using Low-Rank Adaptation (LoRA, rank r=16r=16, α=32\alpha=32, dropout 0.10.1).
      • Task instruction text ii is fed into Q-Former cross-attention to extract instruction-relevant visual features, while questions qq are passed directly to the LLM.
      • UMT-L, Q-Former, linear projection, and LLM LoRA modules are jointly tuned using VTG loss on 2M instruction samples using 8-frame inputs for 3 epochs.
  5. Knowl 5 — Multi-Modal Instruction Dataset for VideoChat2

    model/method

    The instruction-tuning dataset for VideoChat2 contains 2.0M samples curated across 34 public datasets. Every sample is unified into a dictionary structure with visual media paths and QA lists containing task instructions ii, questions qq, and answers aa:

    {
      "video": "path/to/video.mp4",
      "QA": [{
        "i": "Go through the video, taking into account key aspects, and respond to the question.",
        "q": "What color cliff is the hindu temple on?",
        "a": "The Hindu temple in the video is situated on a green cliff."
      }]
    }
    

    Task instructions ii are generated using ChatGPT based on dataset descriptions and task definitions. The 2M samples are divided into 6 functional categories:

    1. Conversation (83.8K83.8\text{K} samples): Multi-turn dialogue from LLaVA, VideoChat, and VideoChatGPT.
    2. Simple Caption (975.5K975.5\text{K} samples): COCO, WebVid, and YouCook2.
    3. Detailed Caption (87.7K87.7\text{K} samples): MiniGPT-4, LLaVA, VideoChat, Paragraph Captioning, TextCaps, TextVR.
    4. VQA (338.4K338.4\text{K} samples): VQAv2, GQA, OK-VQA, A-OKVQA, ViQuAE, OCR-VQA, TextVQA, ST-VQA, DocVQA, TGIF-QA, WebVidQA, Ego4D.
    5. Reasoning (208.4K208.4\text{K} samples): LLaVA-reasoning, CLEVR, VisualMRC, NExT-QA, CLEVRER (QA and Multiple Choice).
    6. Classification (139.9K139.9\text{K} samples): ImageNet, COCO-ITM, Kinetics-710, Something-Something V2.
  6. Knowl 6 — Evaluation Results of MLLMs on MVBench

    empirical result

    Evaluating image-based and video-based MLLMs across the 20 tasks of MVBench reveals that standard MLLMs achieve performance close to a blind text-only baseline, while VideoChat2 achieves superior temporal understanding.

    Model Avg AS AP AA FA UA OE OI OS MD AL ST AC MC MA SC FP CO EN ER CI
    Random 27.3 25.0 25.0 33.3 25.0 25.0 33.3 25.0 33.3 25.0 25.0 25.0 33.3 25.0 33.3 33.3 25.0 33.3 25.0 20.0 30.9
    Image MLLMs (4 frames)
    mPLUG-Owl-I 29.4 25.0 20.0 44.5 27.0 23.5 36.0 24.0 34.0 23.0 24.0 34.5 34.5 22.0 31.5 40.0 24.0 37.0 25.5 21.0 37.0
    LLaMA-Adapter 31.7 23.0 28.0 51.0 30.0 33.0 53.5 32.5 33.5 25.5 21.5 30.5 29.0 22.5 41.5 39.5 25.0 31.5 22.5 28.0 32.0
    BLIP2 31.4 24.5 29.0 33.5 17.0 42.0 51.5 26.0 31.0 25.5 26.0 32.5 25.5 30.0 40.0 42.0 27.0 30.0 26.0 37.0 31.0
    Otter-I 33.5 34.5 32.0 39.5 30.5 38.5 48.5 44.0 29.5 19.0 25.5 55.0 20.0 32.5 28.5 39.0 28.0 27.0 32.0 29.0 36.5
    MiniGPT-4 18.8 16.0 18.0 26.0 21.5 16.0 29.5 25.5 13.0 11.5 12.0 9.5 32.5 15.5 8.0 34.0 26.0 29.5 19.0 9.9 3.0
    InstructBLIP 32.5 20.0 16.5 46.0 24.5 46.0 51.0 26.0 37.5 22.0 23.0 46.5 42.5 26.5 40.5 32.0 25.5 30.0 25.5 30.5 38.0
    LLaVA 36.0 28.0 39.5 63.0 30.5 39.0 53.0 41.0 41.5 23.0 20.5 45.0 34.0 20.5 38.5 47.0 25.0 36.0 27.0 26.5 42.0
    Video MLLMs (16 frames)
    Otter-V 26.8 23.0 23.0 27.5 27.0 29.5 53.0 28.0 33.0 24.5 23.5 27.5 26.0 28.5 18.0 38.5 22.0 22.0 23.5 19.0 19.5
    mPLUG-Owl-V 29.7 22.0 28.0 34.0 29.0 29.0 40.5 27.0 31.5 27.0 23.0 29.0 31.5 27.0 40.0 44.0 24.0 31.0 26.0 20.5 29.5
    VideoChatGPT 32.7 23.5 26.0 62.0 22.5 26.5 54.0 28.0 40.0 23.0 20.0 31.0 30.5 25.5 39.5 48.5 29.0 33.0 29.5 26.0 35.5
    VideoLLaMA 34.1 27.5 25.5 51.0 29.0 39.0 48.0 40.5 38.0 22.5 22.5 43.0 34.0 22.5 32.5 45.5 32.5 40.0 30.0 21.0 37.0
    VideoChat 35.5 33.5 26.5 56.0 33.5 40.5 53.0 40.5 30.0 25.5 27.0 48.5 35.0 20.5 42.5 46.0 26.5 41.0 23.5 23.5 36.0
    VideoChat2text 34.7 24.5 27.0 49.5 27.0 38.0 53.0 28.0 40.0 25.5 27.0 38.5 41.5 27.5 32.5 46.5 26.5 36.0 33.0 32.0 40.0
    VideoChat2 (Vicuna-7B) 51.1 66.0 47.5 83.5 49.5 60.0 58.0 71.5 42.5 23.0 23.0 88.5 39.0 42.0 58.5 44.0 49.0 36.5 35.0 40.5 65.5
    GPT-4V 43.5 55.5 63.5 72.0 46.5 73.5 18.5 59.0 29.5 12.0 40.5 83.5 39.0 12.0 22.5 45.0 47.5 52.0 31.0 59.0 11.0
    VideoChat2 (Mistral-7B) 60.4 75.5 58.0 83.5 50.5 60.5 87.5 74.5 45.0 47.5 44.0 82.5 37.0 64.5 87.5 51.0 66.5 47.0 35.0 37.0 72.5

    Task abbreviations: AS: Action Sequence, AP: Action Prediction, AA: Action Antonym, FA: Fine-grained Action, UA: Unexpected Action, OE: Object Existence, OI: Object Interaction, OS: Object Shuffle, MD: Moving Direction, AL: Action Localization, ST: Scene Transition, AC: Action Count, MC: Moving Count, MA: Moving Attribute, SC: State Change, FP: Fine-grained Pose, CO: Character Order, EN: Egocentric Navigation, ER: Episodic Reasoning, CI: Counterfactual Inference.

    Key findings:

    • VideoChat2 (Vicuna-7B) reaches 51.1%51.1\% accuracy, outperforming VideoChat (35.5%35.5\%) by 15.6%15.6\%.
    • VideoChat2 (Mistral-7B) achieves 60.4%60.4\%, outperforming GPT-4V (43.5%43.5\%) by 16.9%16.9\% overall.
  7. Knowl 7 — VideoChat2 Zero-Shot Video QA and Dialogue Performance

    empirical result

    VideoChat2 establishes state-of-the-art results across standard zero-shot video QA benchmarks and open-ended dialogue evaluations:

    1. Zero-Shot Open-Ended Video QA (Accuracy %\% and GPT-based Score out of 5):

      • MSVD-QA: VideoChat2 achieves 70.0%70.0\% accuracy / 3.93.9 score (vs. VideoChatGPT's 64.9%/3.364.9\% / 3.3, VideoChat's 56.3%/2.856.3\% / 2.8, VideoLLaMA's 51.6%/2.551.6\% / 2.5).
      • MSRVTT-QA: VideoChat2 achieves 54.1%54.1\% accuracy / 3.33.3 score (vs. VideoChatGPT's 49.3%/2.849.3\% / 2.8, VideoChat's 45.0%/2.545.0\% / 2.5).
      • ActivityNet-QA: VideoChat2 achieves 49.1%49.1\% accuracy / 3.33.3 score (vs. VideoChatGPT's 35.2%/2.735.2\% / 2.7, VideoChat's 26.5%/2.226.5\% / 2.2).
    2. Video-ChatGPT Conversation Benchmark (Evaluated by GPT-3.5 across 5 dimensions on a 1–5 scale):

      • Correctness of Information: VideoChat2 scores 3.023.02 (vs. VideoChatGPT 2.402.40, VideoChat 2.232.23).
      • Detail Orientation: VideoChat2 scores 2.882.88 (vs. VideoChatGPT 2.522.52, VideoChat 2.502.50).
      • Contextual Understanding: VideoChat2 scores 3.513.51 (vs. VideoChatGPT 2.622.62, VideoChat 2.532.53).
      • Temporal Understanding: VideoChat2 scores 2.662.66 (vs. VideoChatGPT 1.981.98, VideoChat 1.941.94).
      • Consistency: VideoChat2 scores 2.812.81 (vs. VideoChatGPT 2.372.37, VideoChat 2.242.24).
      • Overall Average: VideoChat2 scores 2.982.98 (vs. VideoChatGPT 2.382.38, VideoChat 2.292.29).
    3. Long-Form Reasoning Benchmarks:

      • EgoSchema Fullset (Zero-Shot): VideoChat2 (Mistral) attains 54.4%54.4\% accuracy using 16 input frames (surpassing InternVideo's 32.1%32.1\% with 90 frames and mPLUG-Owl's 31.1%31.1\% with 5 frames).
      • IntentQA: VideoChat2 (Mistral) achieves 83.4%83.4\% total accuracy on the test set, exceeding the human baseline (78.5%78.5\%).
  8. Knowl 8 — Architectural and Training Strategy Ablations of VideoChat2

    empirical result

    Controlled ablations on MVBench isolate the contribution of vision foundations, LLM backbones, progressive unfreezing, and instruction data compositions:

    1. Visual Encoder vs. LLM Backbone:

      • Replacing EVA-CLIP-g (42.4%42.4\%) with UMT-L (48.6%48.6\%) under frozen LoRA raises MVBench accuracy by +6.2%+6.2\%.
      • Adding LoRA tuning boosts UMT-L + Vicuna-7B v0 from 48.6%48.6\% to 51.1%51.1\% (+2.5%+2.5\%).
      • Scaling the LLM from Vicuna-7B v0 (51.1%51.1\%) to Vicuna-13B v0 (51.4%51.4\%) or Vicuna-13B v1.5 (51.6%51.6\%) yields marginal improvements (≤0.5%≤ 0.5\%), indicating that temporal perception on MVBench is bottlenecked by the visual foundation model rather than LLM parameter count.
    2. Progressive Unfreezing Across Training Stages:

      • Tuning only the linear projection layer (freezing UMT-L and Q-Former) yields 38.5%38.5\%.
      • Unfreezing Q-Former in Stage 2 and Stage 3 increases accuracy to 47.0%47.0\% (+8.5%+8.5\%).
      • Unfreezing both UMT-L and Q-Former in Stage 2 and Stage 3 yields 51.1%51.1\% (+12.6%+12.6\% over the frozen baseline).
    3. Instruction Tuning Data Modality:

      • Training with 1.1M Image instruction samples alone yields 42.1%42.1\%.
      • Training with 0.9M Video instruction samples alone yields 50.5%50.5\% (+8.4%+8.4\% over image-only).
      • Co-training on all 2.0M Image + Video instructions achieves the peak accuracy of 51.1%51.1\%.
  9. Knowl 9 — Frame Count versus Spatial Resolution Trade-Off in Video MLLMs

    empirical result

    Varying visual input sampling parameters during VideoChat2 testing on MVBench demonstrates that temporal capacity depends on temporal frame density rather than spatial resolution:

    Training Input Testing Input MVBench Accuracy (%)
    8×224×2248 \times 224 \times 224 8×224×2248 \times 224 \times 224 50.6
    8×224×2248 \times 224 \times 224 8×384×3848 \times 384 \times 384 49.9 (−0.7-0.7)
    8×224×2248 \times 224 \times 224 16×224×22416 \times 224 \times 224 51.1 (+0.5+0.5)
    8×224×2248 \times 224 \times 224 32×224×22432 \times 224 \times 224 51.1 (+0.5+0.5)
    8×224×2248 \times 224 \times 224 64×224×22464 \times 224 \times 224 51.0 (+0.4+0.4)
    16×224×22416 \times 224 \times 224 16×224×22416 \times 224 \times 224 51.0 (+0.4+0.4)

    Increasing image resolution from 224×224224 \times 224 to 384×384384 \times 384 decreases accuracy by 0.7%0.7\%, whereas increasing the sampled temporal frames from 8 to 16 or 32 frames improves accuracy by +0.5%+0.5\%, confirming that MVBench tasks evaluate temporal dynamics rather than fine spatial pixel details.

  10. Knowl 10 — Failure Modes of Video MLLMs on Spatial Grounding and Counting Tasks

    limitation

    Across both image and video MLLMs evaluated on MVBench, models exhibit severe performance degradations on spatial-temporal localization, object counting, and character order recognition:

    1. Localization and Trajectory Tracking:

      • On Moving Direction (identifying 2D trajectories), VideoChat2 achieves 23.0%23.0\% accuracy, falling below random guess (25.0%25.0\%) and the blind text-only baseline VideoChat2text (25.5%25.5\%).
      • On Action Localization (identifying whether actions occur at the start, middle, end, or throughout), VideoChat2 achieves 23.0%23.0\%, underperforming text-only VideoChat2text (27.0%27.0\%).
    2. Discrete Counting:

      • On Action Count, VideoChat2 attains 39.0%39.0\% (compared to 41.5%41.5\% for text-only VideoChat2text).
    3. Character Sequence Ordering:

      • On Character Order, VideoChat2 achieves 36.5%36.5\%, which is comparable to the text-only baseline (36.0%36.0\%).

    These failure modes stem from the absence of explicit bounding-box grounding and temporal coordinate supervision during instruction tuning, causing vision-language models to struggle with fine-grained spatial and numeric tracking in dynamic videos.

Coverage note — None was omitted. All major contributions—including MVBench task taxonomy and automated curation pipeline, prompt formatting, VideoChat2 progressive training framework, 2M instruction dataset, benchmark evaluation results, zero-shot video QA performance, ablation studies, and stated task limitations—have been systematically captured.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning. ArXiv, abs/2204.14198, 2022. 1, 2
  2. 2.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. ArXiv, abs/2308.12966, 2023. 11
  3. 3.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In ICCV, 2021. 6, 8
  4. 4.Ali Furkan Biten, Ruben Pérez Tito, Andrés Mafla, Lluís Gomez, Marc¸al Rusiñol, Ernest Valveny, C. V. Jawahar, and Dimosthenis Karatzas. Scene text visual question answering. In ICCV, 2019. 6
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In NeurIPS, 2020. 2
  6. 6.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR, 2021. 6
  7. 7.David L. Chen and William B. Dolan. Collecting highly parallel data for paraphrase evaluation. In ACL, 2011. 2
  8. 8.Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny. Minigpt-v2: large language model as a unified interface for vision-language multi-task learning. ArXiv, abs/2310.09478, 2023. 11
  9. 9.Ke Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. Shikra: Unleashing multimodal llm’s referential dialogue magic. ArXiv, abs/2306.15195, 2023. 11
  10. 10.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Díaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. Palm: Scaling language modeling with pathways. JMLR, 2022. 1, 2
  11. 11.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Albert Li, Pascale Fung, and Steven C. H. Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning. In NeurIPS, 2023. 2, 6, 7, 8
  12. 12.Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In NeurIPS, 2022. 9
  13. 13.Pradipto Das, Chenliang Xu, Richard F. Doell, and Jason J. Corso. A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching. In CVPR, 2013. 6
  14. 14.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. 6
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. ArXiv, abs/1810.04805, 2018. 1, 2, 6
  16. 16.Danny Driess, F. Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Ho Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Peter R. Florence. Palm-e: An embodied multimodal language model. In ICML, 2023. 1, 2
  17. 17.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, and Rongrong Ji. Mme: A comprehensive evaluation benchmark for multimodal large language models. ArXiv, abs/2306.13394, 2023. 1, 2, 3
  18. 18.Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, and Zicheng Liu. Violet: End-to-end video-language transformers with masked visual-token modeling. ArXiv, abs/2111.12681, 2021. 10
  19. 19.Difei Gao, Luowei Zhou, Lei Ji, Linchao Zhu, Yezhou Yang, and Mike Zheng Shou. Mist : Multi-modal iterative spatial-temporal transformer for long-form video question answering. In CVPR, 2022. 10
  20. 20.J. Gao, Chen Sun, Zhenheng Yang, and Ramakant Nevatia. Tall: Temporal activity localization via language query. In ICCV, 2017. 3, 12
  21. 21.Tao Gong, Chengqi Lyu, Shilong Zhang, Yudong Wang, Miao Zheng, Qianmengke Zhao, Kuikun Liu, Wenwei Zhang, Ping Luo, and Kai Chen. Multimodal-gpt: A vision and language model for dialogue with humans. ArXiv, abs/2305.04790, 2023. 2
  22. 22.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fründ, Peter Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, and Roland Memisevic. The “something something” video database for learning and evaluating visual common sense. In ICCV, 2017. 2, 6, 9
  23. 23.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In CVPR, 2017. 2, 6
  24. 24.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Z. Xu, Chen Zhao, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Christian Fuegen, Abrham Gebreselasie, Cristina Gonzalez, James M. Hillis, Xuhua Huang, Yifei Huang, Wenqi Jia, Weslie Khoo, Jachym Kolar, Satwik Kottur, Anurag Kumar, Federico Landini, Chao Li, Yanghao Li, Zhenqiang Li, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Yuchen Wang, Xindi Wu, Takuma Yagi, Yunyi Zhu, Pablo Arbelaez, David J. Crandall, Dima Damen, Giovanni Maria Farinella, Bernard Ghanem, Vamsi Krishna Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Kitani, Haizhou Li, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato, Jianbo Shi, Mike Zheng Shou, Antonio Torralba, Lorenzo Torresani, Mingfei Yan, and Jitendra Malik. Ego4d: Around the world in 3,000 hours of egocentric video. In CVPR, 2022. 6
  25. 25.J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In ICLR, 2021. 6, 8, 9
  26. 26.Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Xia Song, and Furu Wei. Language is not all you need: Aligning perception with language models. ArXiv, abs/2302.14045, 2023. 1
  27. 27.Drew A. Hudson and Christopher D. Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In CVPR, 2019. 6
  28. 28.Y. Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In CVPR, 2017. 6
  29. 29.Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L’elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothee Lacroix, and William El Sayed. Mistral 7b. ArXiv, abs/2310.06825, 2023. 7, 10
  30. 30.Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross B. Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In CVPR, 2017. 4, 6
  31. 31.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Apostol Natsev, Mustafa Suleyman, and Andrew Zisserman. The kinetics human action video dataset. ArXiv, abs/1705.06950, 2017. 2
  32. 32.Jacob Krantz, Erik Wijmans, Arjun Majumdar, Dhruv Batra, and Stefan Lee. Beyond the nav-graph: Vision-and-language navigation in continuous environments. In ECCV, 2020. 3, 12
  33. 33.Jonathan Krause, Justin Johnson, Ranjay Krishna, and Li Fei-Fei. A hierarchical approach for generating descriptive image paragraphs. In CVPR, 2017. 6
  34. 34.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV, 2017. 6
  35. 35.Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L. Berg. Tvqa: Localized, compositional video question answering. In EMNLP, 2018. 3, 7, 10, 12
  36. 36.Paul Lerner, Olivier Ferret, Camille Guinaudeau, Hervé Le Borgne, Romaric Besançon, José G. Moreno, and José Lovón-Melgarejo. Viquae, a dataset for knowledge-based visual question answering about named entities. In SIGIR, 2022. 6
  37. 37.Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan. Seed-bench: Benchmarking multimodal llms with generative comprehension. ArXiv, abs/2307.16125, 2023. 2, 9
  38. 38.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. Otter: A multi-modal model with in-context instruction tuning. ArXiv, abs/2305.03726, 2023. 7
  39. 39.Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2022. 1, 6, 7
  40. 40.Jiapeng Li, Ping Wei, Wenjuan Han, and Lifeng Fan. Intentqa: Context-aware video intent reasoning. 2023. 7, 10
  41. 41.Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Limin Wang, and Y. Qiao. Uniformerv2: Spatiotemporal learning by arming image vits with video uniformer. ArXiv, abs/2211.09552, 2022. 6, 10
  42. 42.Kunchang Li, Yinan He, Yi Wang, Yizhuo Li, Wen Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. ArXiv, abs/2305.06355, 2023. 1, 2, 5, 6, 7, 8, 9, 10, 11
  43. 43.Kunchang Li, Yali Wang, Yizhuo Li, Yi Wang, Yinan He, Limin Wang, and Yu Qiao. Unmasked teacher: Towards training-efficient video foundation models. In ICCV, 2023. 2, 6, 8, 9, 10, 12
  44. 44.Lei Li, Yuwei Yin, Shicheng Li, Liang Chen, Peiyi Wang, Shuhuai Ren, Mukai Li, Yazheng Yang, Jingjing Xu, Xu Sun, Lingpeng Kong, and Qi Liu. M3it: A large-scale dataset towards multi-modal multilingual instruction tuning. ArXiv, abs/2306.04387, 2023. 5
  45. 45.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji rong Wen. Evaluating object hallucination in large vision-language models. ArXiv, abs/2305.10355, 2023. 1
  46. 46.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 6, 8
  47. 47.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023. 1, 2, 6, 7, 10
  48. 48.Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot. Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding. TPAMI, 2020. 3, 12
  49. 49.Yuanzhan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. Mmbench: Is your multi-modal model an all-around player? ArXiv, abs/2307.06281, 2023. 1, 2, 3, 5, 8, 9
  50. 50.Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Ming-Hui Qiu, Pengcheng Lu, Tao Wang, and Zhongyu Wei. Valley: Video assistant with large language model enhanced ability. ArXiv, abs/2306.07207, 2023. 2
  51. 51.Muhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. ArXiv, abs/2306.05424, 2023. 2, 6, 7, 8, 10, 11
  52. 52.Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A diagnostic benchmark for very long-form video language understanding. ArXiv, abs/2308.09126, 2023. 7, 10
  53. 53.Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. In CVPR, 2019. 2, 6
  54. 54.Minesh Mathew, Dimosthenis Karatzas, R. Manmatha, and C. V. Jawahar. Docvqa: A dataset for vqa on document images. In WACV, 2021. 6
  55. 55.Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. Ocr-vqa: Visual question answering by reading text in images. In ICDAR, 2019. 6
  56. 56.Mathew Monfort and SouYoung Jin. Spoken moments: Learning joint audio-visual representations from video descriptions. In CVPR, 2021. 7
  57. 57.Mathew Monfort, Bolei Zhou, Sarah Adel Bargal, Alex Andonian, Tom Yan, Kandan Ramakrishnan, Lisa M. Brown, Quanfu Fan, Dan Gutfreund, Carl Vondrick, and Aude Oliva. Moments in time dataset: One million videos for event understanding. TPAMI, 2020. 3, 12
  58. 58.OpenAI. Chatgpt. https://openai.com/blog/chatgpt/, 2023. 1, 4, 5, 8, 10
  59. 59.OpenAI. Gpt-4v(ision) system card. https://api.semanticscholar.org/CorpusID:263218031, 2023. 1, 7
  60. 60.Vicente Ordonez, Girish Kulkarni, and Tamara Berg. Im2text: Describing images using 1 million captioned photographs. In NeurIPS, 2011. 6
  61. 61.Viorica Patraucean, Lucas Smaira, Ankush Gupta, Adria Recasens Continente, Larisa Markeeva, Dylan, Banarse, Mateusz Malinowski, Yezhou Yang, Carl Doersch, Tatiana Matejovicova, Yury Sulsky, Antoine, Miech, Skanda Koppula, Alexander Frechette, Hanna Klimczak, R. Koster, Junlin Zhang, Stephanie, Winkler, Yusuf Aytar, Simon Osindero, Dima Damen, Andrew Zisserman, and Joao Carreira. Perception test : A diagnostic benchmark for multimodal models. In NeurIPS, 2023. 2, 3, 12
  62. 62.Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In ICCV, 2015. 2
  63. 63.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 2020. 2
  64. 64.Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-okvqa: A benchmark for visual question answering using world knowledge. In ECCV, 2022. 6
  65. 65.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL, 2018. 6
  66. 66.Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh. Textcaps: a dataset for image captioning with reading comprehension. In ECCV, 2020. 6
  67. 67.Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In CVPR, 2019. 6
  68. 68.Quan Sun, Yuxin Fang, Ledell Yu Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. ArXiv, abs/2303.15389, 2023. 8
  69. 69.Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida. Visualmrc: Machine reading comprehension on document images. In AAAI, 2021. 6
  70. 70.InternLM Team. Internlm: A multilingual language model with progressively enhanced capabilities. https://github.com/InternLM/InternLM, 2023. 2
  71. 71.Vicuna Team. Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality. https://vicuna.lmsys.org/, 2023. 1, 6, 8
  72. 72.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. ArXiv, abs/2302.13971, 2023. 1, 7, 8
  73. 73.Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony S. Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel M. Kloumann, A. V. Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, R. Subramanian, Xia Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zhengxu Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models. ArXiv, abs/2307.09288, 2023. 2, 8
  74. 74.Alex Jinpeng Wang, Yixiao Ge, Rui Yan, Yuying Ge, Xudong Lin, Guanyu Cai, Jianping Wu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. All in one: Exploring unified video-language pre-training. In CVPR, 2023. 10
  75. 75.Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. Temporal segment networks: Towards good practices for deep action recognition. In ECCV, 2016. 9
  76. 76.Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao. Videomae v2: Scaling video masked autoencoders with dual masking. In CVPR, 2023. 9
  77. 77.Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, Sen Xing, Guo Chen, Junting Pan, Jiashuo Yu, Yali Wang, Limin Wang, and Yu Qiao. Internvideo: General video foundation models via generative and discriminative learning. ArXiv, abs/2212.03191, 2022. 10
  78. 78.Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Jian Ma, Xinyuan Chen, Yaohui Wang, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, and Y. Qiao. Internvid: A large-scale video-text dataset for multimodal understanding and generation. ArXiv, 2023. 6
  79. 79.Zhenhailong Wang, Ansel Blume, Sha Li, Genglin Liu, Jaemin Cho, Zineng Tang, Mohit Bansal, and Heng Ji. Paxion: Patching action knowledge in video-language foundation models. In NeurIPS, 2023. 3, 9, 12
  80. 80.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. Finetuned language models are zero-shot learners. In ICLR, 2021. 2
  81. 81.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. In NeurIPS, 2022. 10
  82. 82.Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B. Tenenbaum, and Chuang Gan. Star: A benchmark for situated reasoning in real-world videos. In NeurIPS, 2021. 3, 4, 7, 10, 12
  83. 83.Weijia Wu, Yuzhong Zhao, Zhuangzi Li, Jiahong Li, Hong Zhou, Mike Zheng Shou, and Xiang Bai. A large cross-modal video retrieval dataset with reading comprehension. ArXiv, abs/2305.03347, 2023. 6
  84. 84.Junbin Xiao, Xindi Shang, Angela Yao, and Tat seng Chua. Next-qa: Next phase of question-answering to explaining temporal actions. In CVPR, 2021. 2, 6, 7, 10
  85. 85.Junbin Xiao, Angela Yao, Zhiyuan Liu, Yicong Li, Wei Ji, and Tat seng Chua. Video as conditional graph hierarchy for multi-granular question answering. In AAAI, 2022. 10
  86. 86.Junbin Xiao, Pan Zhou, Tat seng Chua, and Shuicheng Yan. Video graph transformer for video question answering. In ECCV, 2022. 10
  87. 87.Binzhu Xie, Sicheng Zhang, Zitang Zhou, Bo Li, Yuanhan Zhang, Jack Hessel, Jingkang Yang, and Ziwei Liu. Funqa: Towards surprising video comprehension. ArXiv, abs/2306.14899, 2023. 2, 3, 12
  88. 88.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video question answering via gradually refined attention over appearance and motion. In ICME, 2017. 2, 7, 8
  89. 89.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In CVPR, 2016. 2
  90. 90.Peng Xu, Wenqi Shao, Kaipeng Zhang, Peng Gao, Shuo Liu, Meng Lei, Fanqing Meng, Siyuan Huang, Yu Jiao Qiao, and Ping Luo. Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models. ArXiv, abs/2306.09265, 2023. 1, 2
  91. 91.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Just ask: Learning to answer questions from millions of narrated videos. In ICCV, 2021. 6
  92. 92.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Zero-shot video question answering via frozen bidirectional language models. In NeurIPS, 2022. 10, 11
  93. 93.Qinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu, Qi Qian, Ji Zhang, and Fei Huang. Hitea: Hierarchical temporal-aware video-language pre-training. In ICCV, 2023. 10
  94. 94.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yi Zhou, Junyan Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qiang Qi, Ji Zhang, and Feiyan Huang. mplug-owl: Modularization empowers large language models with multimodality. ArXiv, abs/2304.14178, 2023. 2, 7, 10
  95. 95.Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B. Tenenbaum. Clevrer: Collision events for video representation and reasoning. In ICLR, 2020. 3, 6, 9, 12
  96. 96.Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal. Self-chained image-language model for video localization and question answering. In NeurIPS, 2023. 10
  97. 97.Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities. ArXiv, abs/2308.02490, 2023. 1, 2
  98. 98.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. Activitynet-qa: A dataset for understanding complex web videos via question answering. In AAAI, 2019. 2, 8
  99. 99.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. Activitynet-qa: A dataset for understanding complex web videos via question answering. In AAAI, 2019. 7
  100. 100.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, P. Zhang, Yuxiao Dong, and Jie Tang. Glm-130b: An open bilingual pre-trained model. In ICLR, 2022. 2
  101. 101.Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding. ArXiv, abs/2306.02858, 2023. 2, 5, 7
  102. 102.Hongjie Zhang, Yi Liu, Lu Dong, Yifei Huang, Zhen-Hua Ling, Yali Wang, Limin Wang, and Yu Qiao. Movqa: A benchmark of versatile question-answering for long-form movie understanding. ArXiv, abs/2312.04817, 2023. 3, 12
  103. 103.Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Jiao Qiao. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. ArXiv, abs/2303.16199, 2023. 7
  104. 104.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. ArXiv, abs/2304.10592, 2023. 1, 2, 6, 7, 8

Citation

MLA
Li, K., et al. “MVBench: A Comprehensive Multi-modal Video Understanding Benchmark”. arXiv, 2023, http://arxiv.org/abs/2311.17005v4.
APA
Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Liu, Y., Wang, Z., Xu, J., Chen, G., Luo, P., Wang, L., & Qiao, Y. (2023). MVBench: A Comprehensive Multi-modal Video Understanding Benchmark. arXiv. http://arxiv.org/abs/2311.17005v4
Chicago
Li, K., Y. Wang, Y. He, et al. 2023. “MVBench: A Comprehensive Multi-modal Video Understanding Benchmark”. arXiv. http://arxiv.org/abs/2311.17005v4.
Harvard
Li, K. et al. (2023) “MVBench: A Comprehensive Multi-modal Video Understanding Benchmark”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.17005v4.
Vancouver
1. Li K, Wang Y, He Y, et al (2023) MVBench: A Comprehensive Multi-modal Video Understanding Benchmark. arXiv

BibTeX

@article{li2023mvbench,
  title = {MVBench: A Comprehensive Multi-modal Video Understanding Benchmark},
  author = {Li, Kunchang and Wang, Yali and He, Yinan and Li, Yizhuo and Wang, Yi and Liu, Yi and Wang, Zun and Xu, Jilan and Chen, Guo and Luo, Ping and Wang, Limin and Qiao, Yu},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.17005v4},
  eprint = {2311.17005}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/