ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos
Jr-Jen ChenYu-Chien LiaoHsi-Che LinYu-Chu YuYen-Chun ChenYu-Chiang Frank Wang
Presents a benchmark and cost-effective data generation pipeline to evaluate and improve multimodal language models on complex cause-and-effect temporal reasoning where questions and answers span non-overlapping video segments.
Modern multimodal artificial intelligence systems have made rapid advances in language and vision understanding, yet they continue to struggle with deeper temporal and causal reasoning in videos. Real-world applications—including autonomous robotics, medical video analysis, and legal investigations—require systems to connect cause-and-effect relationships where an event and its underlying cause or consequence happen at different times. Existing evaluation suites largely test whether a model can match text to a specific video segment, failing to assess how well artificial intelligence tracks sequential dependencies across time.
The article introduces and evaluates REXTIME, a benchmark suite specifically designed to measure how effectively multimodal models perform reasoning across temporally separated video events. Its primary objective is to quantitatively evaluate and enhance model capabilities in answering questions where the visual evidence and the queried event do not occur simultaneously.
To construct this benchmark efficiently, the authors developed a semi-automated generation pipeline that extracts time-aligned video events from public datasets, categorizes their relationships (sequential, cause-and-effect, or means-to-an-end), and generates multiple-choice question-and-answer pairs with targeted model self-verification. The final evaluation set contains 921 validation and 2,143 test questions rigorously verified by human annotators, alongside an unverified training set of 9,695 machine-generated samples. The methodology also introduces a metric to measure the time overlap between questions and answers, ensuring the questions genuinely require across-time comprehension.
The evaluation reveals several critical findings. First, even the most capable proprietary systems lag significantly behind human ability: the top-performing model, GPT-4o, achieved 73.7% question-answering accuracy compared to 88.0% human accuracy, representing a substantial 14.3% performance gap. Second, frontier models fail severely at pinpointing the exact video segment containing the answer, achieving less than half the localization accuracy of human annotators. Third, open-source models perform poorly without customization, scoring around 36% to 40% accuracy in zero-shot settings. However, fine-tuning open-source models on the 9,695 machine-generated training samples dramatically improved their performance, boosting one baseline model from 36.3% to 58.2% accuracy while cutting overall data generation costs by 55% compared to fully manual annotation.
These findings demonstrate that current commercial and research artificial intelligence cannot be fully trusted for autonomous decision-making in time-sensitive video workflows without human oversight. Relying on current models for surveillance, safety monitoring, or legal review carries a notable risk of misidentifying causes and effects across continuous footage. The results clearly show that high performance on traditional visual question answering does not translate to genuine temporal understanding.
Organizations developing or deploying video-based artificial intelligence should adopt across-time benchmarks like REXTIME to evaluate vendor models before production deployment. In addition, engineering teams should incorporate targeted temporal fine-tuning data, using automated generation pipelines to improve model reasoning cost-effectively before investing in expensive manual annotations.
Confidence in these findings is high, supported by multi-annotator human baselines and consistent performance gaps across multiple leading proprietary models. Nevertheless, users should consider key limitations: proprietary model evaluations were conducted on a subset of 300 samples due to access and cost constraints, and the automated training data remains unverified by humans, which may introduce minor label noise during model training.
- Paper: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark, Kunchang Li et al. (2023). Provides the foundational benchmark (MVBench) and automated generation methodology for evaluating dynamic temporal tasks in multi-modal video understanding.
- Paper: Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models, Muhammad Maaz et al. (2023). Establishes standard paradigms for instruction-tuning vision-language models on video QA and detailed temporal conversation.
- Paper: Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning, Juncheng Li et al. (2022). Introduces structured compositional temporal grounding across events and actions, establishing the foundational problem formulation for aligning language to video segments.
- Paper: Dense-Captioning Events in Videos, Ranjay Krishna et al. (2017). Pioneers dense event captioning and temporal event proposal in long untrimmed videos that underpin modern video event reasoning.
- Paper: The “Something Something” Video Database for Learning and Evaluating Visual Common Sense, Raghav Goyal et al. (2017). Introduces fine-grained physical and causal interaction tracking in video, providing essential background for video-based common-sense reasoning.
- Paper: TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning, Xiangyu Zeng 0004 et al. (2025). Extends temporal reasoning in video MLLMs by introducing grounded tuning and temporal position encoding to precisely localize temporally separated events.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). Expands comprehensive evaluation of multimodal LLMs across diverse video durations, subtitles, and modalities in video analysis.
- Paper: LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding, Haoning Wu et al. (2024). Evaluates long-context multi-frame temporal referring and reasoning over extended hour-long interleaved video-subtitle sequences.
- Paper: HourVideo: 1-Hour Video-Language Understanding, Keshigeyan Chandrasegaran et al. (2024). Pushes temporal and causal reasoning evaluation to long-horizon, hour-long egocentric video streams requiring extensive multi-event synthesis.
- Paper: TRACE: Temporal Grounding Video LLM via Causal Event Modeling, Yongxin Guo 0001 et al. (2025). Develops a causal event modeling framework that directly addresses temporal localization and grounding deficiencies in video LLMs.
- Paper: Adaptive Keyframe Sampling for Long Video Understanding, Xi Tang et al. (2025). Applies adaptive keyframe sampling to capture query-relevant temporal evidence across long video timelines effectively.
- Paper: M-LLM Based Video Frame Selection for Efficient Video Understanding, Kai Hu 0010 et al. (2025). Enhances long-video question-answering efficiency through query-guided frame selection to pinpoint relevant timestamps across time.
- Paper: Grounded Question-Answering in Long Egocentric Videos, Shangzhe Di et al. (2024). Unifies temporal moment localization with question answering to ground distant evidence in continuous first-person video.
