Video Question Answering: Datasets, Algorithms and Challenges

Yaoyao ZhongWei JiJunbin XiaoYicong LiWeihong DengTat-Seng Chua

article2022EMNLP119 citations

Presents a comprehensive taxonomy of video question answering that classifies benchmarks by modality and reasoning difficulty while systematically evaluating architectural techniques from spatio-temporal attention to neuro-symbolic reasoning.

Listen

Video question answering systems enable artificial intelligence to interact with dynamic visual environments by answering natural language questions about video content. While early efforts focused primarily on simple visual recognition, real-world deployment requires systems that can perform complex temporal, causal, and multi-modal reasoning. The article systematically reviews the field by categorizing existing datasets and algorithms, analyzing recent technical advances, and outlining key directions for future development.

The article establishes a structured taxonomy covering video-only, multi-modal, and knowledge-based tasks, while also distinguishing simple fact-based questions from complex inference-based questions. The authors evaluate foundational modeling techniques, including recurrent memory networks, attention mechanisms, graph neural networks, neural modular units, neural-symbolic systems, and large-scale pre-trained Transformer architectures across standard industry benchmarks.

The comparative analysis reveals several core findings regarding current model capabilities. Large-scale cross-modal pre-training paired with fine-tuning achieves top performance on factoid benchmarks, reaching 47.9% on MSVD-QA and 43.9% on MSRVTT-QA. However, these systems struggle with dynamic relational reasoning, showing notable performance drops when tasked with complex causal questions. On benchmarks like TGIF-QA, high performance metrics mask dataset limitations such as language bias and shallow visual cues rather than genuine reasoning. On more challenging real-world reasoning benchmarks like NExT-QA, the best automated systems reach roughly 52% to 54% accuracy, lagging significantly behind the human benchmark of 88.4%. Graph neural networks, modular networks, and hierarchical representations provide better relational reasoning capabilities and interpretability, though neural-symbolic approaches remain restricted to synthetic environments.

These findings indicate that while pre-training delivers strong baseline perceptual skills, current systems remain unreliable for autonomous, safety-critical, or high-stakes reasoning tasks due to limited explainability and vulnerability to spurious data correlations. Strategic investments must balance the computational costs of massive pre-training against the need for transparent, structured reasoning architectures that can leverage external commonsense knowledge.

Organizations developing or deploying video question answering systems should transition evaluation frameworks toward causal and temporal benchmarks like NExT-QA and diagnostic consistency tests like AGQA 2.0. Engineering efforts should prioritize integrating graph-structured or modular reasoning into pre-trained visual backbones while incorporating domain knowledge bases. Further research is required to reduce the compute requirements of video pre-training and establish rigorous validation standards for model robustness and interpretability before implementing these systems in production environments.

Cover for Video Question Answering: Datasets, Algorithms and Challenges

Abstract

This survey aims to organize the recent advances in video question answering (VideoQA) and point towards future directions. We firstly categorize the datasets into: 1) normal VideoQA, multi-modal VideoQA and knowledge-based VideoQA, according to the modalities invoked in the question-answer pairs, and 2) factoid VideoQA and inference VideoQA, according to the technical challenges in comprehending the questions and deriving the correct answers. We then summarize the VideoQA techniques, including those mainly designed for Factoid QA (such as the early spatio-temporal attention-based methods and the recent Transformer-based ones) and those targeted at explicit relation and logic inference (such as neural modular networks, neural symbolic methods, and graph-structured methods). Aside from the backbone techniques, we also delve into specific models and derive some common and useful insights either for video modeling, question answering, or for cross-modal correspondence learning. Finally, we present the research trends of studying beyond factoid VideoQA to inference VideoQA, as well as towards the robustness and interpretability. Additionally, we maintain a repository, https://github.com/VRU-NExT/VideoQA, to keep trace of the latest VideoQA papers, datasets, and their open-source implementations if available. With these efforts, we strongly hope this survey could shed light on the follow-up VideoQA research.

Table of Contents

  • 1 Introduction
  • 2 VideoQA Task and Datasets
  • 2.1 Problem Formulation
  • 2.2 Evaluation Metrics
  • 2.3 Datasets
  • 2.4 Main Framework
  • 2.5 Challenges and Meaningful Insights
  • 3 Algorithms
  • 3.1 Methods
  • 3.2 Performance Analysis
  • 4 Future Direction
  • 5 Conclusion
  • Acknowledgements
  • Limitations
  • References
  • A Appendix: Details of VideoQA datasets in the literature
  • B Appendix: Timeline of VideoQA techniques

Knowls

  1. Knowl 1 — VideoQA has two complementary taxonomies

    definition

    Video question answering (VideoQA) can be classified along two independent axes. By the information modalities needed, normal VideoQA relies on video content; multimodal VideoQA also uses resources such as subtitles, transcripts, or movie plots; and knowledge-based VideoQA requires external or commonsense knowledge. Unlike multimodal VideoQA, which may provide a paired text source for a question, knowledge-based VideoQA typically draws on a knowledge base available across the dataset. By reasoning demand, factoid VideoQA asks directly about visible facts—such as objects, attributes, or locations—and mainly tests question understanding and visual recognition, whereas inference VideoQA requires reasoning over relationships among visual facts, especially temporal and causal relationships. A dataset can be described using both axes.

  2. Knowl 2 — VideoQA methods are organized into five principal families

    model/method

    The survey groups VideoQA methods into memory networks, Transformers, graph neural networks, modular networks, and neural-symbolic methods, alongside early attention-based and task-specific approaches. Memory networks store sequential video or language information for retrieval, particularly useful for long movie and television narratives; some models use separate appearance and motion memory or heterogeneous memory to combine these signals. Transformers model long-range dependencies and are often used with cross-modal pre-training. Graph methods represent and propagate information about entities, segments, or relations, with some approaches organizing graph reasoning hierarchically across object, frame, and clip levels. Modular methods compose or hierarchically stack reusable conditional reasoning units, offering flexibility across questions, although some lack explicit logic or require questions to fit predefined subtasks. Neural-symbolic methods parse questions into programs, represent objects and dynamics, and execute symbolic reasoning; the survey notes their reasoning strengths on synthetic data but difficulty with unconstrained videos and natural-language questions.

  3. Knowl 3 — VideoQA includes choice, classification, generation, and counting settings

    definition

    A VideoQA model receives a video and a question and predicts an answer. In multiple-choice QA, it selects one answer from the candidates supplied with the question. In open-ended QA, the answer may be produced by classifying the video-question pair into a fixed answer vocabulary, generating a text sequence word by word, or—when the question asks how many times an action occurs—regressing an integer count. The survey notes that multiple-choice settings are often used to study inference beyond factoid answering because they do not require natural-language answer generation and its evaluation.

  4. Knowl 4 — A common VideoQA system has four functional components

    model/method

    A common VideoQA architecture comprises a video encoder, a question encoder, cross-modal interaction, and an answer decoder. The video encoder may represent frame appearance and clip motion, and can add object-level visual and semantic features such as categories and attributes; these features are often extracted using pre-trained 2D or 3D networks. The question encoder creates token-level language representations, for example from GloVe or BERT, which may then be processed alongside visual sequences by RNNs, CNNs, or Transformers. Cross-modal interaction combines question and video information. The decoder selects among supplied choices, classifies into a fixed answer set, or generates an answer sequence, depending on the task. Encoders may be pre-trained or fine-tuned end to end.

  5. Knowl 5 — Attention, scale coordination, and hierarchy are complementary design principles

    model/method

    The survey identifies several reusable design insights. Attention can select relevant spatial regions and temporal segments; self-attention models dependencies within a modality, while cross-modal co-attention lets questions guide video selection and video content guide question representations. Multi-granularity ensemble coordinates representations at different scales—for example, words, phrases, and sentences in language, and regions, trajectories, frames, and clips in video. Hierarchical learning differs by progressively aggregating local, low-level information into global, higher-level representations, such as objects into actions and events. Cross-modal pre-training learns transferable visual-language representations from large web-scale data for downstream fine-tuning, while multi-step reasoning and causal discovery are additional useful approaches. The survey emphasizes that these principles can be combined rather than treated as mutually exclusive.

  6. Knowl 6 — Representative datasets instantiate distinct modality and reasoning demands

    data/table

    Dataset design reflects both taxonomy axes and ranges from synthetic scenes to real-world videos. MSVD-QA is a factoid, normal VideoQA dataset with 1.9K videos and 50K questions, using web videos and open-ended answers. TGIF-QA is an inference-oriented normal VideoQA dataset of 71K animated GIFs and 165K questions, with both multiple-choice and open-ended tasks. CLEVRER uses 10K synthetic videos and 305K questions for temporal and causal reasoning, with automatically generated multiple-choice and open-ended questions. TVQA is multimodal: it uses 21K television-show videos and 152K manually annotated multiple-choice questions, drawing on subtitles and visual concepts. Social-IQ is multimodal inference QA with 1.2K web videos and 7.5K manually annotated multiple-choice questions about social intelligence. KnowIT VQA is knowledge-based inference QA about television shows, with 12K videos and 24K manually annotated multiple-choice questions. These examples illustrate how datasets vary not just in scale but in the evidence sources and reasoning abilities they target.

  7. Knowl 7 — Cross-modal pre-trained Transformers lead the surveyed factoid results

    empirical result

    In the survey's 2022 compilation of factoid benchmarks, cross-modal pre-trained Transformer models generally outperform other approaches. VIOLET reports 68.9 accuracy on TGIF-QA Frame-QA, 47.9 on MSVD-QA, and 43.9 on MSRVTT-QA; MERLOT reports 69.5 on TGIF-QA Frame-QA. Among approaches without cross-modal pre-training, graph-structured methods are especially common and show strong results. The survey also reports that hierarchical learning and fine-grained object features often help performance. It identifies iVQA as a promising open-ended benchmark because its design addresses language bias.

  8. Knowl 8 — Inference results expose a gap between TGIF-QA and real-world reasoning

    empirical result

    Graph-structured methods, causal discovery, and hierarchical learning show promise on inference VideoQA, and cross-modal pre-training and fine-tuning also improve reported inference results. However, scores on TGIF-QA are very high: VGT reports 95.0 accuracy on the action task and 97.6 on the transition task. The survey cautions that TGIF-QA may not be challenging enough and may contain substantial language bias. NExT-QA, which tests temporal and causal relations among multiple objects in real-world videos, remains harder: VGT reports 55.0 validation accuracy and 53.7 test accuracy, compared with human validation accuracy of 88.4. The survey therefore regards NExT-QA as a more demanding benchmark for visual reasoning in realistic videos.

  9. Knowl 9 — Multimodal and knowledge-based QA require evidence-source flexibility

    empirical result

    Multimodal and knowledge-based VideoQA require models to locate and reason over heterogeneous sources of information, rather than relying on visual content alone. The survey reports that these tasks, like normal VideoQA, benefit from stronger architectures and larger-scale datasets, but emphasizes modality-shifting ability: a model must use whichever source—such as video, subtitles, or knowledge—is relevant to a question. This requirement is distinct from simply fusing modalities in a fixed way.

  10. Knowl 10 — The research agenda moves toward reasoning, efficient transfer, and trustworthy models

    limitation

    The survey identifies four unresolved priorities. First, VideoQA should move beyond recognizing objects, attributes, and actions toward temporal and causal reasoning about objects, actions, and events. Second, knowledge should be incorporated to answer questions beyond what is visible, including commonsense and domain-specific questions, with retrieved knowledge diagnosed to support interpretability and trust. Third, cross-modal pre-training needs to become more computationally efficient and better adapted to inference-oriented tasks, since current strengths are clearest on recognition and shallow description. Fourth, interpretability, robustness, and generalization remain open problems: it is not yet clear how advanced pre-trained models work, when they fail, or how reliably they generalize.

Coverage note — The exhaustive dataset inventory and detailed evaluation-metric definitions are omitted: the dataset knowl gives representative coverage of the taxonomy, while the full inventory and metric specifications serve mainly as reference material rather than additional central synthesis.

References

  1. 1.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6077–6086.
  2. 2.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2425–2433.
  3. 3.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1728–1738.
  4. 4.Shyamal Buch, Cristóbal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles. 2022. Revisiting the" video" in video-language understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2917–2927.
  5. 5.Santiago Castro, Mahmoud Azab, Jonathan Stroud, Cristina Noujaim, Ruoyao Wang, Jia Deng, and Rada Mihalcea. 2020. Lifeqa: A real-life dataset for video question answering. In Proceedings of The 12th Language Resources and Evaluation Conference, pages 4352–4358.
  6. 6.Santiago Castro, Naihao Deng, Pingxuan Huang, Mihai Burzo, and Rada Mihalcea. 2022a. In-the-wild video question answering. In Proceedings of the 29th International Conference on Computational Linguistics, pages 5613–5635.
  7. 7.Santiago Castro, Ruoyao Wang, Pingxuan Huang, Ian Stewart, Oana Ignat, Nan Liu, Jonathan Stroud, and Rada Mihalcea. 2022b. Fiber: Fill-in-the-blanks as a challenging video understanding evaluation framework. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2925–2940.
  8. 8.Feilong Chen, Duzhen Zhang, Minglun Han, Xiuyi Chen, Jing Shi, Shuang Xu, and Bo Xu. 2022. Vlp: A survey on vision-language pre-training. arXiv preprint arXiv:2202.09061.
  9. 9.Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua. 2018. Temporally grounding natural sentence in video. In Proceedings of the 2018 conference on empirical methods in natural language processing, pages 162–171.
  10. 10.Long Chen, Hanwang Zhang, Jun Xiao, Liqiang Nie, Jian Shao, Wei Liu, and Tat-Seng Chua. 2017. Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5659–5667.
  11. 11.Zhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong, et al. 2021. Grounding physical concepts of objects and events through dynamic visual reasoning. International Conference on Learning Representations.
  12. 12.Anoop Cherian, Chiori Hori, Tim K Marks, and Jonathan Le Roux. 2022. (2.5+ 1) d spatio-temporal scene graphs for video question answering. In AAAI.
  13. 13.Seongho Choi, Kyoung-Woon On, Yu-Jung Heo, Ahjeong Seo, Youwon Jang, Minsu Lee, and Byoung-Tak Zhang. 2021. Dramaqa: Character-centered video story understanding with hierarchical qa. In AAAI, volume 35.
  14. 14.Anthony Colas, Seokhwan Kim, Franck Dernoncourt, Siddhesh Gupte, Daisy Zhe Wang, and Doo Soon Kim. 2020. Tutorialvqa: Question answering dataset for tutorial videos. In Proceedings of The 12th Language Resources and Evaluation Conference, pages 5450–5455.
  15. 15.Long Hoang Dang, Thao Minh Le, Vuong Le, and Truyen Tran. 2021. Hierarchical object-oriented spatio-temporal reasoning for video question answering. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, pages 636–642.
  16. 16.Mingyu Ding, Zhenfang Chen, Tao Du, Ping Luo, Josh Tenenbaum, and Chuang Gan. 2021. Dynamic visual reasoning by learning differentiable physics models from video and language. Advances in Neural Information Processing Systems, 34.
  17. 17.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations.
  18. 18.Deniz Engin, François Schnitzler, Ngoc QK Duong, and Yannis Avrithis. 2021. On the hidden treasure of dialog in video question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2064–2073.
  19. 19.Alex Falcon, Oswald Lanz, and Giuseppe Serra. 2020. Data augmentation techniques for the video question answering task. In European Conference on Computer Vision, pages 511–525. Springer.
  20. 20.Chenyou Fan. 2019. Egovqa-an egocentric video question answering benchmark dataset. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0.
  21. 21.Chenyou Fan, Xiaofan Zhang, Shu Zhang, et al. 2019. Heterogeneous memory enhanced multimodal attention model for video question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1999–2007.
  22. 22.Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, et al. 2020. Video2commonsense: Generating commonsense descriptions to enrich video captioning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, pages 840–860.
  23. 23.Li Fei-Fei and Ranjay Krishna. 2022. Searching for computer vision north stars. Journal of the American Academy of Arts & Sciences, page 85.
  24. 24.Christiane Fellbaum. 1998. Wordnet. In Wiley Online Library.
  25. 25.Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, et al. 2021. Violet: End-to-end video-language transformers with masked visual-token modeling. arXiv:2111.12681.
  26. 26.Mona Gandhi, Mustafa Omer Gul, Eva Prakash, Madeleine Grunde-McLaughlin, Ranjay Krishna, and Maneesh Agrawala. 2022. Measuring compositional consistency for video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
  27. 27.Difei Gao, Ruiping Wang, Ziyi Bai, and Xilin Chen. 2021. Env-qa: A video question answering benchmark for comprehensive understanding of dynamic environments. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1675–1685.
  28. 28.Jiyang Gao, Runzhou Ge, Kan Chen, and Ram Nevatia. 2018. Motion-appearance co-memory networks for video question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6576–6585.
  29. 29.Noa Garcia and Yuta Nakashima. 2020. Knowledge-based video question answering with unsupervised scene descriptions. In European Conference on Computer Vision, pages 581–598. Springer.
  30. 30.Noa Garcia, Mayu Otani, Chenhui Chu, and Yuta Nakashima. 2020. Knowit vqa: Answering knowledge-based questions about videos. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 10826–10834.
  31. 31.Madeleine Grunde-McLaughlin, Ranjay Krishna, and Maneesh Agrawala. 2021. Agqa: A benchmark for compositional spatio-temporal reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11287–11297.
  32. 32.Pranay Gupta and Manish Gupta. 2022. Newskvqa: Knowledge-aware news video question answering. arXiv preprint arXiv:2202.04015.
  33. 33.Vivek Gupta, Badri N Patro, Hemant Parihar, and Vinay P Namboodiri. 2022. Vquad: Video question answering diagnostic dataset. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 282–291.
  34. 34.Deng Huang, Peihao Chen, Runhao Zeng, Qing Du, Mingkui Tan, and Chuang Gan. 2020. Location-aware graph convolutional networks for video question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11021–11028.
  35. 35.Yunseok Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2019. Video question answering with spatio-temporal reasoning. International Journal of Computer Vision, 127(10):1385–1412.
  36. 36.Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2758–2766.
  37. 37.Jianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao, and Yue Gao. 2020. Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11101–11108.
  38. 38.Pin Jiang and Yahong Han. 2020. Reasoning with heterogeneous graph alignment for video question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11109–11116.
  39. 39.Weike Jin, Zhou Zhao, Yimeng Li, Jie Li, Jun Xiao, and Yueting Zhuang. 2019. Video question answering via knowledge-based progressive spatial-temporal attention network. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 15(2s):1–22.
  40. 40.Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, pages 4171–4186.
  41. 41.Khushboo Khurana and Umesh Deshpande. 2021. Video question-answering techniques, benchmark datasets and evaluation metrics leveraging video captioning: A comprehensive survey. IEEE Access, 9:43799–43823.
  42. 42.Junyeong Kim, Minuk Ma, Kyungsu Kim, Sungjin Kim, et al. 2019. Progressive attention memory network for movie story question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8337–8346.
  43. 43.Junyeong Kim, Minuk Ma, Trung Pham, Kyungsu Kim, and Chang D Yoo. 2020. Modality shifting attention network for multi-modal video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10106–10115.
  44. 44.Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, and Byoung-Tak Zhang. 2017. Deepstory: video story qa by deep embedded memory networks. In Proceedings of the 26th International Joint Conference on Artificial Intelligence, pages 2016–2022.
  45. 45.Nayoung Kim, Seong Jong Ha, and Je-Won Kang. 2021a. Video question answering using language-guided deep compressed-domain video feature. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1708–1717.
  46. 46.Seonhoon Kim, Seohyeong Jeong, Eunbyul Kim, Inho Kang, and Nojun Kwak. 2021b. Self-supervised pre-training and contrastive representation learning for multiple-choice video qa. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 13171–13179.
  47. 47.Thomas N Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. International Conference on Learning Representations.
  48. 48.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123(1):32–73.
  49. 49.Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran. 2020. Hierarchical conditional relation networks for video question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9972–9981.
  50. 50.Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran. 2021. Hierarchical conditional relation networks for multimodal video question answering. International Journal of Computer Vision, 129(11):3027–3050.
  51. 51.Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. 2021. Less is more: Clipbert for video-and-language learning via sparse sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7331–7341.
  52. 52.Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg. 2018. Tvqa: Localized, compositional video question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1369–1379.
  53. 53.Jie Lei, Licheng Yu, Tamara Berg, and Mohit Bansal. 2020. Tvqa+: Spatio-temporal grounding for video question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8211–8225.
  54. 54.Fangtao Li, Ting Bai, Chenyu Cao, Zihe Liu, Chenghao Yan, and Bin Wu. 2021a. Relation-aware hierarchical attention framework for video question answering. In Proceedings of the 2021 International Conference on Multimedia Retrieval, pages 164–172.
  55. 55.Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, and Di Hu. 2022a. Learning to answer questions in dynamic audio-visual scenarios. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19108–19118.
  56. 56.Jiangtong Li, Li Niu, and Liqing Zhang. 2022b. From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21273–21282.
  57. 57.Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. 2020. Hero: Hierarchical encoder for video+ language omni-representation pre-training. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2046–2065.
  58. 58.Linjie Li, Jie Lei, Zhe Gan, and Jingjing Liu. 2021b. Adversarial vqa: A new benchmark for evaluating the robustness of vqa models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2042–2051.
  59. 59.Xiangpeng Li, Jingkuan Song, Lianli Gao, Xianglong Liu, et al. 2019. Beyond rnns: Positional self-attention with co-attention for video question answering. In AAAI, pages 8658–8665.
  60. 60.Yicong Li, Xiang Wang, Junbin Xiao, and Tat-Seng Chua. 2022c. Equivariant and invariant grounding for video question answering. In Proceedings of the 30th ACM International Conference on Multimedia, pages 4714–4722.
  61. 61.Yicong Li, Xiang Wang, Junbin Xiao, Wei Ji, and Tat-Seng Chua. 2022d. Invariant grounding for video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2928–2937.
  62. 62.Yongqi Li, Wenjie Li, and Liqiang Nie. 2022e. Mm-coqa: Conversational question answering over text, tables, and images. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4220–4231.
  63. 63.Fei Liu, Jing Liu, Weining Wang, and Hanqing Lu. 2021a. Hair: Hierarchical visual-semantic relational reasoning for video question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1698–1707.
  64. 64.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. 2021b. Video swin transformer. arXiv preprint arXiv:2106.13230.
  65. 65.Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016. Hierarchical question-image co-attention for visual question answering. Advances in neural information processing systems, 29.
  66. 66.Tegan Maharaj, Nicolas Ballas, Anna Rohrbach, Aaron Courville, and Christopher Pal. 2017. A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6884–6893.
  67. 67.Jianguo Mao, Wenbin Jiang, Xiangdong Wang, Zhifan Feng, Yajuan Lyu, Hong Liu, and Yong Zhu. 2022. Dynamic multistep reasoning based on video scene graph for video question answering. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3894–3904.
  68. 68.Jonghwan Mun, Paul Hongsuck Seo, Ilchae Jung, and Bohyung Han. 2017. Marioqa: Answering questions by watching gameplay videos. In Proceedings of the IEEE International Conference on Computer Vision, pages 2867–2875.
  69. 69.Seil Na, Sangho Lee, Jisung Kim, and Gunhee Kim. 2017. A read-write memory network for movie story understanding. In Proceedings of the IEEE International Conference on Computer Vision, pages 677–685.
  70. 70.Liqiang Nie, Wenjie Wang, Richang Hong, Meng Wang, and Qi Tian. 2019. Multimodal dialog system: Generating responses via adaptive decoders. In Proceedings of the 27th ACM International Conference on Multimedia, pages 1098–1106.
  71. 71.Jungin Park, Jiyoung Lee, and Kwanghoon Sohn. 2021. Bridge to answer: Structure-aware graph interaction network for video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15526–15535.
  72. 72.Devshree Patel, Ratnam Parikh, and Yesha Shastri. 2021. Recent advances in video question answering: A review of datasets and methods. In International Conference on Pattern Recognition, pages 339–356. Springer.
  73. 73.Liang Peng, Shuangji Yang, Yi Bin, and Guoqing Wang. 2021. Progressive graph attention network for video question answering. In Proceedings of the 29th ACM International Conference on Multimedia, pages 2871–2879.
  74. 74.Min Peng, Chongyang Wang, Yuan Gao, Yu Shi, and Xiang-Dong Zhou. 2022. Multilevel hierarchical network with multiscale sampling for video question answering. IJCAI.
  75. 75.AJ Piergiovanni, Kairo Morton, Weicheng Kuo, Michael S Ryoo, and Anelia Angelova. 2022. Video question answering with iterative video-text co-tokenization. European Conference on Computer Vision.
  76. 76.Zi Qian, Xin Wang, Xuguang Duan, Hongyang Chen, and Wenwu Zhu. 2022. Dynamic spatio-temporal modular network for video question answering. In Proceedings of the 30th ACM International Conference on Multimedia, page 4466–4477.
  77. 77.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763.
  78. 78.Arka Sadhu, Kan Chen, and Ram Nevatia. 2021. Video question answering with phrases via semantic roles. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2460–2478.
  79. 79.Ahjeong Seo, Gi-Cheon Kang, Joonhan Park, and Byoung-Tak Zhang. 2021a. Attend what you need: Motion-appearance synergistic networks for video question answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, pages 6167–6177.
  80. 80.Paul Hongsuck Seo, Arsha Nagrani, and Cordelia Schmid. 2021b. Look before you speak: Visually contextualized utterances. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16877–16887.
  81. 81.Xindi Shang, Donglin Di, Junbin Xiao, Yu Cao, Xun Yang, and Tat-Seng Chua. 2019. Annotating objects and relations in user-generated videos. In Proceedings of the 2019 on International Conference on Multimedia Retrieval, pages 279–287.
  82. 82.Xindi Shang, Yicong Li, Junbin Xiao, Wei Ji, and Tat-Seng Chua. 2021. Video visual relation detection via iterative inference. In Proceedings of the 29th ACM International Conference on Multimedia, pages 3654–3663.
  83. 83.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565.
  84. 84.Sasha Sheng, Amanpreet Singh, Vedanuj Goswami, Jose Magana, Tristan Thrush, Wojciech Galuba, Devi Parikh, and Douwe Kiela. 2021. Human-adversarial visual question answering. Advances in Neural Information Processing Systems, 34:20346–20359.
  85. 85.Xiaomeng Song, Yucheng Shi, Xin Chen, and Yahong Han. 2018. Explore multi-step reasoning in video question answering. In Proceedings of the 26th ACM international conference on Multimedia, pages 239–247.
  86. 86.Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al. 2015. End-to-end memory networks. Advances in neural information processing systems, 28.
  87. 87.Guanglu Sun, Lili Liang, Tianlin Li, Bo Yu, et al. 2021. Video question answering: a survey of models and datasets. Mobile Networks and Applications, 26(5):1904–1937.
  88. 88.Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016. Movieqa: Understanding stories in movies through question-answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4631–4640.
  89. 89.Aisha Urooj, Amir Mazaheri, Mubarak Shah, et al. 2020. Mmft-bert: Multimodal fusion transformer with bert encodings for visual question answering. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4648–4660.
  90. 90.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  91. 91.Jianyu Wang, Bingkun Bao, and Changsheng Xu. 2021. Dualvgr: A dual-visual graph reasoning unit for video question answering. IEEE Transactions on Multimedia.
  92. 92.Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan. 2021a. Star: A benchmark for situated reasoning in real-world videos. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  93. 93.Tianran Wu, Noa Garcia, Mayu Otani, Chenhui Chu, Yuta Nakashima, and Haruo Takemura. 2021b. Transferring domain-agnostic knowledge in video question answering. arXiv:2110.13395.
  94. 94.Zhibiao Wu and Martha Stone Palmer. 1994. Verb semantics and lexical selection. In 32nd Annual Meeting of the Association for Computational Linguistics, 27-30 June 1994, New Mexico State University, Las Cruces, New Mexico, USA, Proceedings, pages 133–138.
  95. 95.Junbin Xiao, Xindi Shang, Xun Yang, Sheng Tang, and Tat-Seng Chua. 2020. Visual relation grounding in videos. In European conference on computer vision, pages 447–464. Springer.
  96. 96.Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. 2021. Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9777–9786.
  97. 97.Junbin Xiao, Angela Yao, Zhiyuan Liu, Yicong Li, Wei Ji, and Tat-Seng Chua. 2022a. Video as conditional graph hierarchy for multi-granular question answering. In Proceedings of the 36th AAAI Conference on Artificial Intelligence (AAAI), pages 2804–2812.
  98. 98.Junbin Xiao, Pan Zhou, Tat-Seng Chua, and Shuicheng Yan. 2022b. Video graph transformer for video question answering. European Conference on Computer Vision.
  99. 99.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. 2017. Video question answering via gradually refined attention over appearance and motion. In Proceedings of the 25th ACM international conference on Multimedia, pages 1645–1653.
  100. 100.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296.
  101. 101.Li Xu, He Huang, and Jun Liu. 2021. Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9878–9888.
  102. 102.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. 2021a. Just ask: Learning to answer questions from millions of narrated videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1686–1697.
  103. 103.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. 2022a. Learning to answer visual questions from web videos. IEEE Transactions on Pattern Analysis and Machine Intelligence.
  104. 104.Pinci Yang, Xin Wang, Xuguang Duan, Hong Chen, Runze Hou, Cong Jin, and Wenwu Zhu. 2022b. Avqa: A dataset for audio-visual question answering on videos. In Proceedings of the 30th ACM International Conference on Multimedia, pages 3480–3491.
  105. 105.Xun Yang, Fuli Feng, Wei Ji, Meng Wang, and Tat-Seng Chua. 2021b. Deconfounded video moment retrieval with causal intervention. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1–10.
  106. 106.Zekun Yang, Noa Garcia, Chenhui Chu, Mayu Otani, Yuta Nakashima, and Haruo Takemura. 2020. Bert representations for video question answering. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1556–1565.
  107. 107.Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, et al. 2020. Clevrer: Collision events for video representation and reasoning. In International Conference on Learning Representations.
  108. 108.Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum. 2018. Neural-symbolic vqa: Disentangling reasoning from vision and language understanding. Advances in neural information processing systems, 31.
  109. 109.Weijiang Yu, Haoteng Zheng, Mengfei Li, Lei Ji, Lijun Wu, Nong Xiao, and Nan Duan. 2021. Learning from inside: Self-driven siamese sampling and reasoning for video question answering. Advances in Neural Information Processing Systems, 34.
  110. 110.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. 2019. Activitynet-qa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 9127–9134.
  111. 111.Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee, and Gunhee Kim. 2021. Pano-avqa: Grounded audio-visual question answering on 360deg videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2031–2041.
  112. 112.Amir Zadeh, Michael Chan, Paul Pu Liang, Edmund Tong, and Louis-Philippe Morency. 2019. Social-iq: A question answering benchmark for artificial social intelligence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8807–8817.
  113. 113.Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. 2021. Merlot: Multimodal neural script knowledge models. Advances in Neural Information Processing Systems, 34.
  114. 114.Kuo-Hao Zeng, Tseng-Hung Chen, Ching-Yao Chuang, Yuan-Hong Liao, Juan Carlos Niebles, and Min Sun. 2017. Leveraging video descriptions to learn video question answering. In Thirty-First AAAI Conference on Artificial Intelligence.
  115. 115.Ao Zhang, Yuan Yao, Qianyu Chen, Wei Ji, Zhiyuan Liu, Maosong Sun, and Tat-Seng Chua. 2022. Fine-grained scene graph generation with data transfer. In European conference on computer vision.
  116. 116.Wentian Zhao, Seokhwan Kim, Ning Xu, and Hailin Jin. 2020. Video question answering on screencast tutorials. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, pages 1061–1068.
  117. 117.Zhou Zhao, Jinghao Lin, Xinghua Jiang, Deng Cai, Xiaofei He, and Yueting Zhuang. 2017a. Video question answering via hierarchical dual-level attention network learning. In Proceedings of the 25th ACM international conference on Multimedia, pages 1050–1058.
  118. 118.Zhou Zhao, Qifan Yang, Deng Cai, Xiaofei He, and Yueting Zhuang. 2017b. Video question answering via hierarchical spatio-temporal attention networks. In IJCAI, volume 2, page 8.
  119. 119.Zhou Zhao, Zhu Zhang, Shuwen Xiao, Zhou Yu, Jun Yu, Deng Cai, Fei Wu, and Yueting Zhuang. 2018. Open-ended long-form video question answering via adaptive hierarchical reinforced networks. In IJCAI, volume 2, page 8.
  120. 120.Linchao Zhu, Zhongwen Xu, Yi Yang, and Alexander G Hauptmann. 2017. Uncovering the temporal context for video question answering. International Journal of Computer Vision, 124(3):409–421.
  121. 121.Linchao Zhu and Yi Yang. 2020. Actbert: Learning global-local video-text representations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8746–8755.
  122. 122.Yueting Zhuang, Dejing Xu, Xin Yan, Wenzhuo Cheng, Zhou Zhao, Shiliang Pu, and Jun Xiao. 2020. Multichannel attention refinement for video question answering. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 16(1s):1–23.

Citation

MLA
Zhong, Y., et al. “Video Question Answering: Datasets, Algorithms and Challenges”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 6439–55, https://doi.org/10.18653/v1/2022.emnlp-main.432.
APA
Zhong, Y., Ji, W., Xiao, J., Li, Y., Deng, W., & Chua, T.-S. (2022). Video Question Answering: Datasets, Algorithms and Challenges. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6439–6455. https://doi.org/10.18653/v1/2022.emnlp-main.432
Chicago
Zhong, Y., W. Ji, J. Xiao, Y. Li, W. Deng, and T.-S. Chua. 2022. “Video Question Answering: Datasets, Algorithms and Challenges”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6439–55. https://doi.org/10.18653/v1/2022.emnlp-main.432.
Harvard
Zhong, Y. et al. (2022) “Video Question Answering: Datasets, Algorithms and Challenges”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 6439–6455. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.432.
Vancouver
1. Zhong Y, Ji W, Xiao J, Li Y, Deng W, Chua T-S (2022) Video Question Answering: Datasets, Algorithms and Challenges. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 6439–6455

BibTeX

@inproceedings{zhong-etal-2022-video,
    title = "Video Question Answering: Datasets, Algorithms and Challenges",
    author = "Zhong, Yaoyao  and
      Ji, Wei  and
      Xiao, Junbin  and
      Li, Yicong  and
      Deng, Weihong  and
      Chua, Tat-Seng",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.432/",
    doi = "10.18653/v1/2022.emnlp-main.432",
    pages = "6439--6455"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/