ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos

Jr-Jen ChenYu-Chien LiaoHsi-Che LinYu-Chu YuYen-Chun ChenYu-Chiang Frank Wang

article2024NeurIPS67 citations

Presents a benchmark and cost-effective data generation pipeline to evaluate and improve multimodal language models on complex cause-and-effect temporal reasoning where questions and answers span non-overlapping video segments.

Listen

Modern multimodal artificial intelligence systems have made rapid advances in language and vision understanding, yet they continue to struggle with deeper temporal and causal reasoning in videos. Real-world applications—including autonomous robotics, medical video analysis, and legal investigations—require systems to connect cause-and-effect relationships where an event and its underlying cause or consequence happen at different times. Existing evaluation suites largely test whether a model can match text to a specific video segment, failing to assess how well artificial intelligence tracks sequential dependencies across time.

The article introduces and evaluates REXTIME, a benchmark suite specifically designed to measure how effectively multimodal models perform reasoning across temporally separated video events. Its primary objective is to quantitatively evaluate and enhance model capabilities in answering questions where the visual evidence and the queried event do not occur simultaneously.

To construct this benchmark efficiently, the authors developed a semi-automated generation pipeline that extracts time-aligned video events from public datasets, categorizes their relationships (sequential, cause-and-effect, or means-to-an-end), and generates multiple-choice question-and-answer pairs with targeted model self-verification. The final evaluation set contains 921 validation and 2,143 test questions rigorously verified by human annotators, alongside an unverified training set of 9,695 machine-generated samples. The methodology also introduces a metric to measure the time overlap between questions and answers, ensuring the questions genuinely require across-time comprehension.

The evaluation reveals several critical findings. First, even the most capable proprietary systems lag significantly behind human ability: the top-performing model, GPT-4o, achieved 73.7% question-answering accuracy compared to 88.0% human accuracy, representing a substantial 14.3% performance gap. Second, frontier models fail severely at pinpointing the exact video segment containing the answer, achieving less than half the localization accuracy of human annotators. Third, open-source models perform poorly without customization, scoring around 36% to 40% accuracy in zero-shot settings. However, fine-tuning open-source models on the 9,695 machine-generated training samples dramatically improved their performance, boosting one baseline model from 36.3% to 58.2% accuracy while cutting overall data generation costs by 55% compared to fully manual annotation.

These findings demonstrate that current commercial and research artificial intelligence cannot be fully trusted for autonomous decision-making in time-sensitive video workflows without human oversight. Relying on current models for surveillance, safety monitoring, or legal review carries a notable risk of misidentifying causes and effects across continuous footage. The results clearly show that high performance on traditional visual question answering does not translate to genuine temporal understanding.

Organizations developing or deploying video-based artificial intelligence should adopt across-time benchmarks like REXTIME to evaluate vendor models before production deployment. In addition, engineering teams should incorporate targeted temporal fine-tuning data, using automated generation pipelines to improve model reasoning cost-effectively before investing in expensive manual annotations.

Confidence in these findings is high, supported by multi-annotator human baselines and consistent performance gaps across multiple leading proprietary models. Nevertheless, users should consider key limitations: proprietary model evaluations were conducted on a subset of 300 samples due to access and cost constraints, and the automated training data remains unverified by humans, which may introduce minor label noise during model training.

arXiv: 2406.19392
Cover for ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos

Abstract

We introduce REXTIME, a benchmark designed to rigorously test AI models’ ability to perform temporal reasoning within video events. Specifically, REXTIME focuses on reasoning across time, i.e. human-like understanding when the question and its corresponding answer occur in different video segments. This form of reasoning, requiring advanced understanding of cause-and-effect relationships across video segments, poses significant challenges to even the frontier multimodal large language models. To facilitate this evaluation, we develop an automated pipeline for generating temporal reasoning question-answer pairs, significantly reducing the need for labor-intensive manual annotations. Our benchmark includes 921 carefully vetted validation samples and 2,143 test samples, each manually curated for accuracy and relevance. Evaluation results show that while frontier large language models outperform academic models, they still lag behind human performance by a significant 14.3% accuracy gap. Additionally, our pipeline creates a training dataset of 9,695 machine generated samples without manual effort, which empirical studies suggest can enhance the across-time reasoning via fine-tuning.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Data collection
  • 3.1 Selecting videos to annotate
  • 3.2 Question-answering on two events across time
  • 3.3 Balancing cheap machine generated data and high-quality human annotation
  • 4 Benchmark
  • 4.1 Evaluation metrics
  • 4.2 How far are frontier MLLMs to solving REXTIME?
  • 4.3 Are academic and open source models competitive?
  • 4.4 Dataset statistics
  • 5 Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Reasoning-Across-Time Video Question Answering

    definition

    Reasoning-across-time is defined as a video question answering (QA) setting in which the video segment referred to by the question does not completely overlap with the video segment containing the answer or evidence. The REXTIME benchmark formalizes this task to evaluate multimodal AI models on causal, intentional, and sequential temporal dependencies across distant video events.

    The benchmark comprises:

    • Test split: 2,143 multiple-choice QA pairs, with ground-truth answer event time spans re-annotated and vetted by human judges.
    • Validation split: 921 human-curated and time-annotated multiple-choice QA pairs.
    • Training split: 9,695 unverified, fully machine-generated QA pairs generated without manual labor to support model fine-tuning.

    Each sample is structured as a 4-choice multiple-choice QA task where a model must predict the correct answer option and localize the start and end timestamps of the answer event in the video.

  2. Knowl 2 — Question-Answer Intersection over Union and Certificate Length

    equation

    To quantify the degree of across-time reasoning required by a video QA dataset, two temporal metrics are defined:

    1. Question-Answer Intersection over Union (QA-IoU): Given a question event temporal span TQ=[tQ,start,tQ,end]T_Q = [t_{Q,\text{start}}, t_{Q,\text{end}}] and an answer event temporal span TA=[tA,start,tA,end]T_A = [t_{A,\text{start}}, t_{A,\text{end}}], the QA-IoU measures temporal overlap: QA-IoU=∣TQ∩TA∣∣TQ∪TA∣\text{QA-IoU} = \frac{|T_Q \cap T_A|}{|T_Q \cup T_A|} where ∣⋅∣|\cdot| denotes duration in seconds. A lower mean QA-IoU (QA-mIoU) indicates that the question and answer events are temporally separated, requiring models to reason across distinct segments rather than searching within a single co-located window.

    2. Certificate Length (C.L.): The minimal duration of the video segment required to answer a question, defined as the span from the earliest start timestamp to the latest end timestamp covering both the question and answer: C.L.=max⁡(tQ,end,tA,end)−min⁡(tQ,start,tA,start)\text{C.L.} = \max(t_{Q,\text{end}}, t_{A,\text{end}}) - \min(t_{Q,\text{start}}, t_{A,\text{start}}) A longer Certificate Length signifies that the model must process and reason over a wider temporal duration in the video.

  3. Knowl 3 — Event Relation Taxonomy and Classification Rules for Temporal Video QA

    model/method

    Candidate video event pairs (E1,E2)(E_1, E_2) extracted from caption annotations are categorized into three distinct relation types based on four criteria scored by GPT-4 on an integer scale s∈{0,1,2,3}s \in \{0, 1, 2, 3\}:

    • Directness: Evaluates the directness of the causal link connecting E1E_1 to E2E_2.
    • Necessity: Evaluates whether E2E_2 is an inevitable consequence of E1E_1 (i.e., whether E2E_2 would still occur without E1E_1).
    • Proactiveness: Evaluates whether E1E_1 is performed with deliberate human intent.
    • Purpose: Evaluates whether the intended goal of E1E_1 is achieved in E2E_2.

    Event pairs are classified according to the following deterministic rules:

    1. Sequential: Assigned if Directness+Necessity<4\text{Directness} + \text{Necessity} < 4. Non-consecutive event pairs in this category are pruned to avoid temporal ambiguity in "before / after" queries.
    2. Cause-Effect: Assigned if Directness+Necessity≥4\text{Directness} + \text{Necessity} \ge 4 and Proactiveness+Purpose<5\text{Proactiveness} + \text{Purpose} < 5.
    3. Means-to-an-End: Assigned if Directness+Necessity≥4\text{Directness} + \text{Necessity} \ge 4 and Proactiveness+Purpose≥5\text{Proactiveness} + \text{Purpose} \ge 5.
  4. Knowl 4 — Automated LLM-Assisted Pipeline for Generating Temporal Video QA

    model/method

    The REXTIME data creation pipeline generates video QA pairs across non-overlapping video segments while reducing human labor and lowering generation costs from 300to300 to 135 per 1,000 QA pairs. The pipeline consists of four stages:

    1. Stage I: Event Pair Selection:

      • ActivityNet (dense captions): GPT-4 extracts event pairs with distinct timestamps and potential causal relationships, filtering out semantically duplicate pairs using an LLM-evaluated similarity threshold.
      • QVHighlights (sparse captions): For each annotated pivotal event, the video clip is extended by 10 seconds before and 10 seconds after (10 frames at 2 s/frame). GPT-4V detects preceding causes and succeeding effects from the extended visual frames.
    2. Stage II: Event Relation Categorization:

      • GPT-4 rates the event pairs on Directness, Necessity, Proactiveness, and Purpose to classify them into Sequential, Cause-Effect, or Means-to-an-End relations.
    3. Stage III: Question-Answer Generation:

      • Relation-specific in-context learning prompts instruct an LLM to generate a question, a correct answer, and three negative distractor options.
    4. Stage IV: Verification and Curation:

      • The LLM self-evaluates the logical validity of generated Cause-Effect and Means-to-an-End QA pairs to filter out erroneous outputs.
      • Unverified surplus outputs form the 9,695-sample training set.
      • A subset is vetted by human annotators who check logical consistency, correct option validity, and re-annotate the precise time span of the answer event to yield the validation (921 samples) and test (2,143 samples) splits.
  5. Knowl 5 — Dataset Statistics and Comparison with Temporal Grounding Benchmarks

    data/table

    Compared to prior video QA and moment retrieval benchmarks (Ego4D-NLQ and NExTGQA), REXTIME features a substantially larger average Certificate Length (C.L.) and a significantly lower Question-Answer mean Intersection over Union (QA-mIoU).

    Datasets # of Reasoning Across Time Samples C.L. (s) ↑\uparrow QA-mIoU (%) ↓\downarrow
    Train Val Test
    Ego4D-NLQ 2,212 775 705 5.2 85.5
    NExTGQA – 1,403 2,301 11.7 66.1
    REXTIME 9,695 921 2,143 66.0 15.5

    The average Certificate Length in REXTIME is 66.0 seconds compared to 5.2 seconds in Ego4D-NLQ and 11.7 seconds in NExTGQA. The QA-mIoU is 15.5%, compared to 85.5% in Ego4D-NLQ and 66.1% in NExTGQA, demonstrating that REXTIME requires reasoning over distinct, non-overlapping visual event intervals.

  6. Knowl 6 — Performance of Frontier Multimodal LLMs and Humans on REXTIME

    data/table

    Evaluation of proprietary frontier Multimodal Large Language Models (MLLMs) and human annotators on a mini-test split of 300 samples (100 randomly chosen per event relation category: Sequential, Cause-Effect, Means-to-an-End). Metrics evaluate multiple-choice VQA accuracy (Accuracy (%)), moment localization mean Intersection over Union (mIoU), Recall@1 at IoU thresholds of 0.3 and 0.5 (R@1), and joint accuracy requiring both correct answer selection and moment grounding with IoU≥0.5\text{IoU} \ge 0.5 (Accuracy (%) @ IoU ≥\ge 0.5).

    Models Moment Localization VQA
    mIoU R@1 (IoU=0.3) R@1 (IoU=0.5) Accuracy (%) Accuracy (%) @ IoU ≥\ge 0.5
    Human 61.11 74.30 62.85 87.98 58.51
    GPT-4o 36.28 45.33 34.00 73.67 28.67
    Claude3-Opus 23.61 30.67 17.67 68.67 13.67
    Gemini-1.5-Pro 28.43 35.67 25.00 68.00 18.33
    GPT-4V 26.74 33.33 22.00 63.33 16.67
    Reka-Core 27.95 36.33 24.00 59.67 17.00

    Frontier MLLMs exhibit a substantial gap compared to human performance: the top-performing model, GPT-4o, achieves 73.67% VQA accuracy (14.31% below human accuracy of 87.98%) and an mIoU of 36.28 (24.83 below human mIoU of 61.11).

  7. Knowl 7 — Zero-Shot Evaluation of Open-Source Video Models on REXTIME

    empirical result

    Zero-shot evaluation of open-source moment retrieval models (UniVTG, CG-DETR) and grounding video LLMs (VTimeLLM, TimeChat, LITA) on the REXTIME test split (2,143 samples) demonstrates weak zero-shot reasoning-across-time capabilities:

    • UniVTG: mIoU = 28.17, R@1 (IoU=0.3) = 41.34, R@1 (IoU=0.5) = 26.88 (VQA not supported).
    • CG-DETR: mIoU = 23.87, R@1 (IoU=0.3) = 31.31, R@1 (IoU=0.5) = 16.67 (VQA not supported).
    • VTimeLLM: mIoU = 20.14, R@1 (IoU=0.3) = 28.84, R@1 (IoU=0.5) = 17.41, VQA Accuracy = 36.16%.
    • TimeChat: mIoU = 11.65, R@1 (IoU=0.3) = 14.42, R@1 (IoU=0.5) = 7.61, VQA Accuracy = 40.04%.
    • LITA: mIoU = 21.49, R@1 (IoU=0.3) = 29.49, R@1 (IoU=0.5) = 16.29, VQA Accuracy = 34.44%.

    The best zero-shot VQA accuracy among open-source models is 40.04% (TimeChat), performing only modestly above the 25.0% random baseline for 4-option multiple-choice questions.

  8. Knowl 8 — Test Performance of Open-Source Video Models After Fine-Tuning on REXTIME

    data/table

    Fine-tuning open-source models on the 9,695 machine-generated training samples from the REXTIME automated pipeline substantially improves both answer moment localization and temporal VQA accuracy on the 2,143-sample test split.

    Models Moment Localization VQA
    mIoU R@1 (IoU=0.3) R@1 (IoU=0.5) Accuracy (%) Accuracy (%) @ IoU ≥\ge 0.5
    UniVTG 34.63 53.48 34.53 – –
    CG-DETR 26.53 39.71 22.73 – –
    VTimeLLM 29.92 43.69 26.13 57.58 17.13
    TimeChat 26.29 40.13 21.42 49.46 10.92

    Fine-tuning elevates VTimeLLM's VQA accuracy from 36.16% to 57.58% (an increase of 21.42 percentage points), approaching frontier model zero-shot performance (e.g., Reka-Core at 59.67%), and boosts its moment localization mIoU from 20.14 to 29.92. TimeChat's VQA accuracy improves from 40.04% to 49.46% and its mIoU from 11.65 to 26.29.

  9. Knowl 9 — Language-Only Shortcut Assessment via BlindQA

    empirical result

    To determine whether multiple-choice QA pairs in REXTIME can be solved through unimodal linguistic biases without visual video comprehension, a BlindQA experiment was conducted where the model receives only the question and the four candidate options without the video.

    Using GPT-4o:

    • Full Multimodal VQA Accuracy: 73.67%
    • BlindQA (Text-Only) Accuracy: 29.67%

    The performance drop of 44.00 percentage points down to near-random performance (25.0% for 4 options) confirms that REXTIME questions cannot be solved via language shortcuts alone and strictly require visual temporal reasoning across the video content.

Coverage note — Detailed prompt texts and specific dollar cost breakdowns per API call were omitted as standard implementation specifics, as all core methods, classification taxonomy, benchmark statistics, and empirical results are fully captured.

References

  1. 1.The claude 3 model family: Opus, sonnet, haiku. Technical report, Anthropic, 2024. 1, 6, 7
  2. 2.Gpt-4 system card. Technical report, OpenAI, 2024. 1, 6, 7
  3. 3.Reka core, flash, and edge: A series of powerful multimodal language models. Technical report, Reka, 2024. 6, 7
  4. 4.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 1, 3, 6, 7
  5. 5.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691, 2022. 1
  6. 6.Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. Localizing moments in video with natural language. In ICCV, 2017. 3
  7. 7.Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In CVPR, 2015. 1, 4
  8. 8.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 3
  9. 9.Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua. Temporally grounding natural sentence in video. In EMNLP, 2018. 3
  10. 10.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning. In NeurIPS, 2024. 3
  11. 11.Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. A survey on in-context learning. arXiv preprint arXiv:2301.00234, 2022. 6
  12. 12.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In ICCV, 2019. 3
  13. 13.Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. Tall: Temporal activity localization via language query. In ICCV, 2017. 1, 3
  14. 14.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In CVPR, 2022. 3, 8, 9
  15. 15.Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. Localizing moments in video with temporal language. In EMNLP, 2018. 3
  16. 16.Bin Huang, Xin Wang, Hong Chen, Zihan Song, and Wenwu Zhu. Vtimellm: Empower llm to grasp video moments. In CVPR, 2024. 3, 7, 8
  17. 17.De-An Huang, Shijia Liao, Subhashree Radhakrishnan, Hongxu Yin, Pavlo Molchanov, Zhiding Yu, and Jan Kautz. Lita: Language instructed temporal-localization assistant. arXiv preprint arXiv:2403.19046, 2024. 3, 7, 8
  18. 18.Jinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon, and Kwanghoon Sohn. Knowing where to focus: Event-aware transformer for video grounding. In ICCV, 2023. 3
  19. 19.Emre Kıcıman, Robert Ness, Amit Sharma, and Chenhao Tan. Causal reasoning and large language models: Opening a new frontier for causality. arXiv preprint arXiv:2305.00050, 2023. 1
  20. 20.Jie Lei, Tamara L Berg, and Mohit Bansal. Detecting moments and highlights in videos via natural language queries. In NeurIPS, 2021. 1, 3, 4, 6
  21. 21.KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023. 3
  22. 22.Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. Hero: Hierarchical encoder for video+ language omni-representation pre-training. In EMNLP, 2020. 1
  23. 23.Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick, Difei Gao, Alex Jinpeng Wang, Rui Yan, and Mike Zheng Shou. Univtg: Towards unified video-language temporal grounding. In ICCV, 2023. 3, 7, 8
  24. 24.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2024. 3
  25. 25.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424, 2023. 3
  26. 26.Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A diagnostic benchmark for very long-form video language understanding. In NeurIPS, 2024. 3, 8
  27. 27.WonJun Moon, Sangeek Hyun, SuBeen Lee, and Jae-Pil Heo. Correlation-guided query-dependency calibration in video representation learning for temporal grounding. arXiv preprint arXiv:2311.08835, 2023. 3, 7, 8
  28. 28.WonJun Moon, Sangeek Hyun, SangUk Park, Dongchan Park, and Jae-Pil Heo. Query-dependent video representation for moment retrieval and highlight detection. In CVPR, 2023. 3
  29. 29.Abhishek Padalkar, Acorn Pooley, Ajinkya Jain, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anikait Singh, Anthony Brohan, et al. Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864, 2023. 1
  30. 30.Piotr Padlewski, Max Bain, Matthew Henderson, Zhongkai Zhu, Nishant Relan, Hai Pham, Donovan Ong, Kaloyan Aleksiev, Aitor Ormazabal, Samuel Phua, et al. Vibe-eval: A hard evaluation suite for measuring progress of multimodal language models. arXiv preprint arXiv:2405.02287, 2024. 2
  31. 31.Long Qian, Juncheng Li, Yu Wu, Yaobo Ye, Hao Fei, Tat-Seng Chua, Yueting Zhuang, and Siliang Tang. Momentor: Advancing video large language model with fine-grained temporal reasoning. In ICML, 2024. 3
  32. 32.Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, and Lu Hou. Timechat: A time-sensitive multimodal large language model for long video understanding. In CVPR, 2024. 3, 7, 8
  33. 33.Yale Song, Jordi Vallmitjana, Amanda Stent, and Alejandro Jaimes. Tvsum: Summarizing web videos using titles. In CVPR, 2015. 1
  34. 34.Andrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure, Rin Metcalf, Walter Talbott, Natalie Mackraz, R Devon Hjelm, and Alexander T Toshev. Large language models as generalizable policies for embodied tasks. In CoRL, 2024. 1
  35. 35.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. 1, 3, 6, 7
  36. 36.Yueqian Wang, Xiaojun Meng, Jianxin Liang, Yuxuan Wang, Qun Liu, and Dongyan Zhao. Hawkeye: Training video-text llms for grounding text in videos. arXiv preprint arXiv:2403.10228, 2024. 3
  37. 37.Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. Next-qa: Next phase of question-answering to explaining temporal actions. In CVPR, 2021. 1, 3, 8
  38. 38.Junbin Xiao, Angela Yao, Yicong Li, and Tat Seng Chua. Can i trust your answer? visually grounded video question answering. In CVPR, 2024. 3, 6, 8, 9
  39. 39.Huijuan Xu, Kun He, Bryan A Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko. Multilevel language and vision integration for text-to-clip retrieval. In AAAI, 2019. 3
  40. 40.Antoine Yang, Arsha Nagrani, Ivan Laptev, Josef Sivic, and Cordelia Schmid. Vidchapters-7m: Video chapters at scale. 2023. 1
  41. 41.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. mplug-owl: Modularization empowers large language models with multimodality. In CVPR, 2024. 3
  42. 42.Yitian Yuan, Tao Mei, and Wenwu Zhu. To find where you talk: Temporal sentence localization in video with attention based location regression. In AAAI, 2019. 3
  43. 43.Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding. In EMNLP, 2023. 3
  44. 44.Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou. Span-based localizing network for natural language video localization. In ACL, 2020. 3
  45. 45.Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. In ICLR, 2024. 3
  46. 46.Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo. Learning 2d temporal adjacent networks for moment localization with natural language. In AAAI, 2020. 3
  47. 47.Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et al. Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment. In ICLR, 2024. 3
  48. 48.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. In ICLR, 2024. 3
  49. 49.Ming Zhu, Aman Ahuja, Da-Cheng Juan, Wei Wei, and Chandan K Reddy. Question answering with long multiple-span answers. In EMNLP, 2020. 1

Citation

MLA
Chen, J.-J., et al. “ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos”. arXiv, 2024, http://arxiv.org/abs/2406.19392v2.
APA
Chen, J.-J., Liao, Y.-C., Lin, H.-C., Yu, Y.-C., Chen, Y.-C., & Wang, Y.-C. F. (2024). ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos. arXiv. http://arxiv.org/abs/2406.19392v2
Chicago
Chen, J.-J., Y.-C. Liao, H.-C. Lin, Y.-C. Yu, Y.-C. Chen, and Y.-C. F. Wang. 2024. “ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos”. arXiv. http://arxiv.org/abs/2406.19392v2.
Harvard
Chen, J.-J. et al. (2024) “ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2406.19392v2.
Vancouver
1. Chen J-J, Liao Y-C, Lin H-C, Yu Y-C, Chen Y-C, Wang Y-CF (2024) ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos. arXiv

BibTeX

@article{chen2024rextime,
  title = {ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos},
  author = {Chen, Jr-Jen and Liao, Yu-Chien and Lin, Hsi-Che and Yu, Yu-Chu and Chen, Yen-Chun and Wang, Yu-Chiang Frank},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2406.19392v2},
  eprint = {2406.19392}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission