Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Chiara PlizzariAlessio TonioniYongqin XianAce KulshresthaFederico Tombari

article2025CVPR46 citations

Introduces EgoTempo, a benchmark designed to expose static frame and language shortcuts in egocentric video question answering and rigorously evaluate whether multimodal models can reason across full video sequences.

Listen

Artificial intelligence systems increasingly process video to assist humans in complex real-world tasks, making first-person, egocentric video comprehension vital. Egocentric video captures continuous, close-up interactions with tools and environments, requiring models to track fine-grained actions and objects over time. However, evaluating whether multi-modal large language models truly reason over time remains challenging because existing benchmarks often allow models to succeed through commonsense language guessing or static, single-frame visual clues rather than genuine temporal reasoning.

The article aims to rigorously evaluate and expose the temporal understanding limitations of modern multi-modal large language models in egocentric environments. To achieve this, it introduces EgoTempo, a new open-ended video question-answering benchmark designed specifically to require full-video comprehension across diverse first-person scenarios.

The authors constructed EgoTempo by extracting video clips averaging 45 seconds across 40 real-world scenario categories from the Ego4D dataset. Using a semi-automated pipeline combining vision-language models and manual filtering, they generated 500 validated question-answer pairs divided across 10 temporal capabilities covering actions and objects. They evaluated 13 leading commercial and open-source models, including Gemini, GPT-4o, Claude, and open-source architectures, assessing accuracy across text-only inputs, single frames, and multi-frame inputs sampled at varying rates.

The key findings reveal significant limitations in current model capabilities. First, the article demonstrates that existing egocentric benchmarks are heavily solvable without video: leading models achieve 41% to 51% accuracy on existing benchmarks using just a single frame, and up to 58% using only text, whereas on EgoTempo single-frame accuracy drops to just 9.1%. Second, current state-of-the-art models perform poorly on temporal reasoning even when provided with the full video: the highest-performing model, GPT-4o, achieved only 44.4% accuracy with 64 frames, and Gemini reached 39.1% at one frame per second. Third, while increasing frame counts improves performance from 9.1% to 39.1% (a 4.3-fold increase), accuracy quickly plateaus around 40%, showing that simply feeding more video frames into expanded context windows is insufficient. Fourth, tasks involving sequential ordering and counting proved especially difficult, showing little to no improvement from additional visual frames. Finally, a human baseline study achieved 63.2% average accuracy, outperforming the best model by roughly 24 percentage points and confirming a substantial machine-human performance gap.

These results imply that enterprise and research stakeholders should not assume that large context windows or high single-image benchmark scores translate to reliable video comprehension in time-critical, first-person applications such as robotics, manufacturing, or healthcare assistance. Relying on current models for sequential action verification or object tracking carries substantial operational risk due to frequent misinterpretations of temporal order and repeated actions.

The article recommends that AI development move beyond merely increasing input frame counts or context windows. Technical roadmaps should prioritize architectures and training techniques that explicitly model temporal dependencies, sequence transitions, and dynamic interactions. In addition, practitioners evaluating video models should adopt open-ended temporal benchmarks rather than standard multiple-choice datasets that disguise temporal reasoning failures through commonsense guessing.

Regarding confidence and limitations, the findings are supported by consistent evaluations across 13 distinct models and cross-validated automated grading showing a 96% alignment with human judgment. However, the benchmark scope is limited to 500 curated questions from 40 scenarios with an average duration of 45 seconds. Additionally, lower human performance on certain sequence tasks indicates inherent subjectivity in action granularity, meaning stakeholders should treat current performance numbers as conservative indicators of egocentric video reasoning complexity.

No sufficiently relevant recommendations were found.

Cover for Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Abstract

Understanding fine-grained temporal dynamics is crucial in egocentric videos, where continuous streams capture frequent, close-up interactions with objects. In this work, we bring to light that current egocentric video question-answering datasets often include questions that can be answered using only few frames or commonsense reasoning, without being necessarily grounded in the actual video. Our analysis shows that state-of-the-art Multi-Modal Large Language Models (MLLMs) on these benchmarks achieve remarkably high performance using just text or a single frame as input. To address these limitations, we introduce EgoTempo, a dataset specifically designed to evaluate temporal understanding in the egocentric domain. EgoTempo emphasizes tasks that require integrating information across the entire video, ensuring that models would need to rely on temporal patterns rather than static cues or pre-existing knowledge. Extensive experiments on EgoTempo show that current MLLMs still fall short in temporal reasoning on egocentric videos, and thus we hope EgoTempo will catalyze new research in the field and inspire models that better capture the complexity of temporal dynamics. Dataset and code are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 EgoTempo Dataset
  • 3.1 Capability Taxonomy
  • 3.2 Data Curation Pipeline
  • 4 Experiments
  • 4.1 Experiment Setup
  • 4.2 Results
  • 4.3 Ablations
  • 4.4 Qualitative Results
  • 5 Conclusion
  • References
  • A Analysis on Evaluation
  • A.1 Prompt for OpenQA Evaluation
  • A.2 Gemini vs GPT4 for Evaluation
  • A.3 Error Analysis
  • B CloseQA vs OpenQA
  • B.1 Prompt for Q&A Generation
  • C Human Evaluation
  • D Additional Qualitative Results

Knowls

  1. Knowl 1 — EgoTempo is a manually filtered benchmark for temporal reasoning in egocentric video

    definition

    EgoTempo is a free-form, open-ended video question-answering benchmark built from egocentric videos in Ego4D. It contains 500 question-answer pairs, with 50 pairs for each of 10 capabilities, drawn from 365 annotated clips across 221 source videos and 40 scenarios. The clips range from 3 to 140 seconds, with an average duration of 45 seconds. Questions were retained only when answering required information across multiple frames and the question and answer were consistent with the video; examples answerable from a single frame or lacking temporal complexity were excluded.

  2. Knowl 2 — The benchmark separates temporal understanding into ten action and object capabilities

    definition

    EgoTempo’s action-focused capabilities are: Action Sequence, identifying the order of a person’s actions; Action Counting, counting repetitions of an action; Temporal Event Ordering, identifying actions before or after a specified event; Future Action Prediction, predicting the immediate next action; and Object-Specific Actions, identifying what the person does with a specified object. Its object-focused capabilities are: Object Sequence, ordering objects the person interacts with; Object Counting, counting objects present or interacted with; Spatial Relations, determining spatial relationships between objects or between an object and the camera wearer; Locating Objects, identifying an object’s location at a specified time; and Action-Specific Objects, identifying the object used before, during, or after a specified action.

  3. Knowl 3 — EgoTempo combines narration-based clip assembly, automatic question generation, and video-based manual review

    model/method

    The construction pipeline starts from Ego4D narration sentences NjN_j with timestamps tjt_j. For narration jj in video ii, the paper assigns a temporal window Tj=(tj−βi2α,tj+βi2α)T_j=(t_j-\frac{\beta_i}{2\alpha},t_j+\frac{\beta_i}{2\alpha}), where βi\beta_i is the average interval between consecutive narration timestamps in video ii, and α\alpha is the global average of those intervals across the dataset. Consecutive narration segments, which average about 5 seconds, are grouped into clips until either 120 narrations or a 120-second duration is reached, whichever comes first.

    Gemini 1.5 Pro is used zero-shot in two stages: it first generates a caption from the clip and narrations, then generates questions and answers for the ten capabilities from the caption and video. Human reviewers check each pair against the video for multi-frame dependence and logical consistency, and remove irrelevant, insufficiently temporal, or single-frame-answerable pairs. This process yields the 500 benchmark pairs.

  4. Knowl 4 — Existing egocentric VideoQA datasets permit substantially more text-only or single-frame answering

    empirical result

    In the paper’s cross-dataset comparison using Gemini Flash, EgoTempo requires more video evidence than the compared egocentric datasets. The reported accuracy values are: EgoTaskQA, 57.8% using question text only, 40.6% using one central frame, and 41.3%, 55.7%, and 54.9% at sampling rates of 0.1, 0.5, and 1 frame per second (FPS); EgoSchema, 31.3%, 50.6%, 69.8%, 70.4%, and 70.5%, respectively; EgoThink, 27.6% text-only and 65.5% single-frame; and EgoTempo, 10.0%, 9.2%, 19.0%, 32.7%, and 39.1%, respectively. EgoTempo’s accuracy at 1 FPS is 4.3 times its single-frame accuracy, compared with 1.3 times for EgoTaskQA and 1.4 times for EgoSchema. Thus, in this evaluation, added frames benefit EgoTempo much more, while text-only and central-frame performance are much lower.

  5. Knowl 5 — Open-ended benchmark performance remains low across current video MLLMs

    empirical result

    The paper evaluates open-ended predictions with Gemini 1.5 Pro as an answer judge, comparing question-only, central-frame, and multiple-frame inputs. In the reported multiple-frame configuration, GPT-4o reaches 40.1% average accuracy and Gemini 1.5 Flash reaches 39.1%. The other reported multiple-frame averages are Qwen2-VL-72B 28.4%, LLaVA-OneVision-Qwen2-72B 26.5%, Qwen2-VL-7B 26.1%, LLaVA-OneVision-Qwen2-7B 23.3%, InternLM-XC2.5 21.7%, Qwen2-VL-2B 21.3%, MiniCPM-V2.6 18.6%, LLaVA-NeXT-Video-34B 15.7%, Claude-3.5-Sonnet 13.1%, and VideoLLaVA 12.7%. Random-answer accuracy is 4.9%; Gemini Flash scores 10.0% from question text alone and 9.1% from a single frame. In the Gemini 1 FPS result, category accuracy ranges from 12.0% for action counting to 64.0% for spatial relations, with 24.0% for object sequence. The results show that multiple frames help, but do not bring the tested models to high overall accuracy.

  6. Knowl 6 — More frames raise average accuracy, but gains vary by model and do not remove temporal errors

    empirical result

    In a frame-count ablation, average accuracy for Gemini 1.5 Flash at 1, 4, 8, 16, 32, and 64 frames is 9.1%, 24.3%, 27.9%, 34.9%, 36.1%, and 38.9%. For GPT-4o, the corresponding values are 12.9%, 28.9%, 35.3%, 40.8%, 40.1%, and 44.4%. The largest gains generally occur when moving beyond a single frame; GPT-4o’s average dips slightly from 16 to 32 frames before rising at 64. Gemini’s performance also rises with sampling rate, reaching 39.1% at 1 FPS, but the paper reports that gains begin to saturate above 1 FPS, with performance around 40% even at higher rates. The authors also report that low-frame inputs become less effective as video duration increases, whereas performance is more stable with more frames. More visual context therefore helps, but does not by itself ensure reliable temporal reasoning.

  7. Knowl 7 — Chronological frame order contributes to EgoTempo performance

    empirical result

    Gemini Flash was evaluated with uniform sampling (U), random sampling (R), and uniform sampling followed by chronological shuffling (S). Average accuracies for U, R, and shuffled S are, respectively, 39.1%, 30.5%, and 32.3% at 1 FPS; 34.1%, 28.5%, and 32.0% at 0.5 FPS; and 19.0%, 17.9%, and 16.0% at 0.1 FPS. At 1 FPS, shuffling lowers accuracy by 6.8 percentage points relative to uniform chronological input; at 0.1 FPS, it lowers accuracy by 3.0 points. The authors specifically note performance drops on sequence and temporal-event-ordering tasks when uniform frames are shuffled, supporting the benchmark’s sensitivity to temporal order rather than just the presence of sampled visual content.

  8. Knowl 8 — Question format strongly changes measured accuracy

    empirical result

    The paper compares open-ended QA (OpenQA) with four-option multiple-choice QA (CloseQA) for Gemini Flash. For question-text-only input, average accuracy is 10.0% in OpenQA and 36.2% in CloseQA; for a single frame, it is 9.1% and 43.8%; and for multiple frames sampled at 1 FPS, it is 39.1% and 60.9%. Random-choice accuracy is 4.9% for OpenQA and 25.0% for CloseQA. The CloseQA distractors were generated as three plausible but incorrect options. Because CloseQA substantially raises accuracy even without video, the authors use OpenQA for the main benchmark to reduce the possibility of answering from text cues or option selection alone.

  9. Knowl 9 — Human participants outperform Gemini Flash, but the benchmark remains difficult for people

    empirical result

    Twenty participants answered EgoTempo questions after viewing the corresponding videos. Their average accuracy is 63.2%, compared with 39.1% for Gemini Flash and 4.9% random-answer accuracy, a 24.1-percentage-point human–model gap. Human accuracy is 78.0% on action counting and 76.0% on object counting, but 25.0% on action sequence and 44.9% on object sequence. The authors hypothesize that sequence questions are particularly difficult for people because the granularity used to identify a sequence can be subjective; performance on counting also remains imperfect.

  10. Knowl 10 — The LLM-based answer judge showed 96% agreement with human review in a small audit

    empirical result

    The benchmark evaluation compares each model prediction with its ground-truth answer using Gemini 1.5 Pro as an LLM judge; the judge returns a correct/incorrect decision and a score from 0 to 5. In an audit of 100 question-answer pairs, with 10 pairs from each capability, reviewers identified four cases where Gemini marked an answer incorrect although a human would accept it, corresponding to 96% agreement on that sample. The reported error patterns were predictions containing more specific detail than the reference answer and object descriptions that were understandable but used an imprecise name. This audit supports the evaluation procedure while identifying these sources of disagreement.

Coverage note — Detailed category-by-category model tables, qualitative examples, and the full video-duration curves are omitted because they provide secondary diagnostic detail beyond the benchmark definition, headline comparisons, and temporal-order and frame-count analyses retained here.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716–23736, 2022. 2
  2. 2.Anthropic. Claude-sonnet-3.5, 2024. Accessed: 2024-11-13. 5, 6
  3. 3.Jing Bi, Yunlong Tang, Luchuan Song, Ali Vosoughi, Nguyen Nguyen, and Chenliang Xu. Eagle: Egocentric aggregated language-video engine. arXiv preprint arXiv:2409.17523, 2024. 3
  4. 4.Mu Cai, Reuben Tan, Jianrui Zhang, Bocheng Zou, Kai Zhang, Feng Yao, Fangrui Zhu, Jing Gu, Yiwu Zhong, Yuzhang Shang, et al. Temporalbench: Towards fine-grained temporal understanding for multimodal video models. 11, 12
  5. 5.Mu Cai, Jianwei Yang, Jianfeng Gao, and Yong Jae Lee. Matryoshka multimodal models. arXiv preprint arXiv:2405.17430, 2024. 11, 12
  6. 6.Sijie Cheng, Zhicheng Guo, Jingwen Wu, Kechen Fang, Peng Li, Huaping Liu, and Yang Liu. Egothink: Evaluating first-person perspective thinking capability of vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14291–14302, 2024. 3, 4
  7. 7.Shangzhe Di and Weidi Xie. Grounded question-answering in long egocentric videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12934–12943, 2024. 3, 5, 11
  8. 8.Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Xilin Wei, Songyang Zhang, Haodong Duan, Maosong Cao, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Xinyue Zhang, Wei Li, Jingwen Li, Kai Chen, Conghui He, Xingcheng Zhang, Yu Qiao, Dahua Lin, and Jiaqi Wang. Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model. arXiv preprint arXiv:2401.16420, 2024. 2, 5, 6
  9. 9.Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023. 2
  10. 10.Chenyou Fan. Egovqa-an egocentric video question answering benchmark dataset. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019. 3, 4
  11. 11.Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075, 2024. 3
  12. 12.Rohit Girdhar and Deva Ramanan. Cater: A diagnostic dataset for compositional actions and temporal reasoning. arXiv preprint arXiv:1910.04744, 2019. 3
  13. 13.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18995–19012, 2022. 3, 4
  14. 14.Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, et al. Language is not all you need: Aligning perception with language models. Advances in Neural Information Processing Systems, 36:72096–72109, 2023. 2
  15. 15.Baoxiong Jia, Ting Lei, Song-Chun Zhu, and Siyuan Huang. Egotaskqa: Understanding human tasks in egocentric videos. Advances in Neural Information Processing Systems, 35:3343–3360, 2022. 1, 2, 3, 4
  16. 16.Muhammad Uzair khattak, Muhammad Ferjad Naeem, Jameel Hassan, Naseer Muzzamal, Federico Tombari, Fahad Shahbaz Khan, and Salman Khan. How good is my video lmm? complex video reasoning and robustness evaluation suite for video-lmms. arXiv:2405.03690, 2024. 11
  17. 17.Muhammad Uzair Khattak, Muhammad Ferjad Naeem, Jameel Hassan, Muzammal Naseer, Federico Tombari, Fahad Shahbaz Khan, and Salman Khan. Complex video reasoning and robustness evaluation suite for video-lmms. arXiv preprint arXiv:2405.03690, 2024. 5
  18. 18.Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024. 1, 5, 6
  19. 19.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning, pages 12888–12900. PMLR, 2022. 2
  20. 20.Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22195–22206, 2024. 3
  21. 21.Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122, 2023. 2, 5, 6
  22. 22.Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Z Xu, Difei Gao, Rong-Cheng Tu, Wenzhe Zhao, Weijie Kong, et al. Egocentric video-language pretraining. Advances in Neural Information Processing Systems, 35:7575–7586, 2022. 4
  23. 23.Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, 2024. 2
  24. 24.Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, and Lu Hou. Tempcompass: Do video llms really understand videos? arXiv preprint arXiv: 2403.00476, 2024. 3
  25. 25.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424, 2023. 2, 5
  26. 26.Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A diagnostic benchmark for very long-form video language understanding. Advances in Neural Information Processing Systems, 36:46212–46244, 2023. 1, 2, 3, 4
  27. 27.OpenAI. Chatgpt. https://openai.com/blog/chatgpt/, 2023. Accessed: 2023. 2
  28. 28.OpenAI. Gpt-4 technical report, 2023. 2
  29. 29.OpenAI. Gpt-4v(ision) system card. https://cdn.openai.com/papers/GPTV_System_Card.pdf, 2023. Accessed: 2023. 2, 11
  30. 30.OpenAI. Gpt-4o. https://openai.com/index/hello-gpt-4o/, 2024. Accessed: 2024-11-15. 1, 2, 5, 6
  31. 31.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023. 2
  32. 32.Yusu Qian, Haotian Zhang, Yinfei Yang, and Zhe Gan. How easy is it to fool your multimodal llms? an empirical analysis on deceptive prompts. arXiv preprint arXiv:2402.13220, 2024. 5
  33. 33.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 2
  34. 34.Reuben Tan, Ximeng Sun, Ping Hu, Jui-hsien Wang, Hanieh Deilamsalehy, Bryan A Plummer, Bryan Russell, and Kate Saenko. Koala: Key frame-conditioned long video-llm. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13581–13591, 2024. 5
  35. 35.Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024. 1, 2, 3, 5, 6, 7, 8, 11, 12
  36. 36.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste ´ Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. ` Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 2
  37. 37.Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024. 1, 2, 5, 6
  38. 38.Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9777–9786, 2021. 2
  39. 39.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296, 2016. 3
  40. 40.Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800, 2024. 2, 5, 6
  41. 41.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. Activitynet-qa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 9127–9134, 2019. 3
  42. 42.Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, et al. Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark. arXiv preprint arXiv:2409.02813, 2024. 11, 12
  43. 43.Jiahao Zhang, Frederic Z Zhang, Cristian Rodriguez, Yizhak Ben-Shabat, Anoop Cherian, and Stephen Gould. Temporally grounding instructional diagrams in unconstrained videos. arXiv preprint arXiv:2407.12066, 2024. 3
  44. 44.Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li. Llava-next: A strong zero-shot video understanding model, 2024. 1, 5, 6

Citation

MLA
Plizzari, C., et al. “Omnia De EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos”. arXiv, 2025, http://arxiv.org/abs/2503.13646v1.
APA
Plizzari, C., Tonioni, A., Xian, Y., Kulshrestha, A., & Tombari, F. (2025). Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos. arXiv. http://arxiv.org/abs/2503.13646v1
Chicago
Plizzari, C., A. Tonioni, Y. Xian, A. Kulshrestha, and F. Tombari. 2025. “Omnia De EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos”. arXiv. http://arxiv.org/abs/2503.13646v1.
Harvard
Plizzari, C. et al. (2025) “Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2503.13646v1.
Vancouver
1. Plizzari C, Tonioni A, Xian Y, Kulshrestha A, Tombari F (2025) Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos. arXiv

BibTeX

@article{plizzari2025omnia,
  title = {Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos},
  author = {Plizzari, Chiara and Tonioni, Alessio and Xian, Yongqin and Kulshrestha, Achin and Tombari, Federico},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2503.13646v1},
  eprint = {2503.13646}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/