Video Question Answering: Datasets, Algorithms and Challenges
Yaoyao ZhongWei JiJunbin XiaoYicong LiWeihong DengTat-Seng Chua
Presents a comprehensive taxonomy of video question answering that classifies benchmarks by modality and reasoning difficulty while systematically evaluating architectural techniques from spatio-temporal attention to neuro-symbolic reasoning.
Video question answering systems enable artificial intelligence to interact with dynamic visual environments by answering natural language questions about video content. While early efforts focused primarily on simple visual recognition, real-world deployment requires systems that can perform complex temporal, causal, and multi-modal reasoning. The article systematically reviews the field by categorizing existing datasets and algorithms, analyzing recent technical advances, and outlining key directions for future development.
The article establishes a structured taxonomy covering video-only, multi-modal, and knowledge-based tasks, while also distinguishing simple fact-based questions from complex inference-based questions. The authors evaluate foundational modeling techniques, including recurrent memory networks, attention mechanisms, graph neural networks, neural modular units, neural-symbolic systems, and large-scale pre-trained Transformer architectures across standard industry benchmarks.
The comparative analysis reveals several core findings regarding current model capabilities. Large-scale cross-modal pre-training paired with fine-tuning achieves top performance on factoid benchmarks, reaching 47.9% on MSVD-QA and 43.9% on MSRVTT-QA. However, these systems struggle with dynamic relational reasoning, showing notable performance drops when tasked with complex causal questions. On benchmarks like TGIF-QA, high performance metrics mask dataset limitations such as language bias and shallow visual cues rather than genuine reasoning. On more challenging real-world reasoning benchmarks like NExT-QA, the best automated systems reach roughly 52% to 54% accuracy, lagging significantly behind the human benchmark of 88.4%. Graph neural networks, modular networks, and hierarchical representations provide better relational reasoning capabilities and interpretability, though neural-symbolic approaches remain restricted to synthetic environments.
These findings indicate that while pre-training delivers strong baseline perceptual skills, current systems remain unreliable for autonomous, safety-critical, or high-stakes reasoning tasks due to limited explainability and vulnerability to spurious data correlations. Strategic investments must balance the computational costs of massive pre-training against the need for transparent, structured reasoning architectures that can leverage external commonsense knowledge.
Organizations developing or deploying video question answering systems should transition evaluation frameworks toward causal and temporal benchmarks like NExT-QA and diagnostic consistency tests like AGQA 2.0. Engineering efforts should prioritize integrating graph-structured or modular reasoning into pre-trained visual backbones while incorporating domain knowledge bases. Further research is required to reduce the compute requirements of video pre-training and establish rigorous validation standards for model robustness and interpretability before implementing these systems in production environments.
- Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). This foundational image-VQA paper establishes the question–visual input–answer task and benchmark framing that the survey extends from still images to video.
- Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). Its taxonomy of multimodal representation, alignment, and fusion provides core concepts for understanding the survey’s treatment of video–question correspondence.
- Paper: Stacked Attention Networks for Image Question Answering, Zichao Yang et al. (2015). Its progressive visual-attention approach supplies an early QA mechanism that helps explain the spatio-temporal attention methods reviewed in the survey.
- Paper: Hierarchical Question-Image Co-Attention for Visual Question Answering, Jiasen Lu et al. (2016). Its joint attention to question words and visual regions is a useful precursor to the cross-modal attention methods the survey discusses for video QA.
- Paper: MSR-VTT: A Large Video Description Dataset for Bridging Video and Language, Jun Xu et al. (2016). MSR-VTT provides a major video–language dataset and benchmark context for understanding the video QA datasets and evaluations surveyed.
- Paper: Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models, Muhammad Maaz et al. (2024). Video-ChatGPT carries the survey’s video-QA methods into open-ended dialogue, testing whether large vision-language models can give detailed answers about video.
- Paper: Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization, Yang Jin et al. (2024). Video-LaVIT extends video QA into unified video-language pretraining, applying the survey’s multimodal and temporal reasoning themes to a broader model framework.
- Paper: LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding, Haoning Wu et al. (2024). LongVideoBench pushes video QA toward hour-long, subtitle-interleaved inputs, extending the survey’s discussion of temporal understanding into long-context evaluation.
- Paper: HourVideo: 1-Hour Video-Language Understanding, Keshigeyan Chandrasegaran et al. (2024). HourVideo continues video-language evaluation at hour-long durations, asking whether systems can synthesize events and reason over timelines beyond the survey’s short-clip focus.
- Paper: ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos, Jr-Jen Chen et al. (2024). ReXTime advances the survey’s call for inference beyond factoid QA by benchmarking answers that require connecting events separated in time.
- Paper: Streaming Video Question-Answering with In-context Video KV-Cache Retrieval, Shangzhe Di et al. (2025). ReKV extends offline video question answering to continuous streams, using query-relevant cache retrieval to address long-video memory and efficiency limits.
- Paper: Grounded Question-Answering in Long Egocentric Videos, Shangzhe Di et al. (2024). GroundVQA continues video QA into long egocentric recordings by jointly locating the relevant moment and generating an answer.
- Paper: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark, Kunchang Li et al. (2023). MVBench extends video-understanding evaluation across temporal perception and cognition tasks, providing a broader benchmark for capabilities adjacent to those surveyed.
