Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
Chiara PlizzariAlessio TonioniYongqin XianAce KulshresthaFederico Tombari
Introduces EgoTempo, a benchmark designed to expose static frame and language shortcuts in egocentric video question answering and rigorously evaluate whether multimodal models can reason across full video sequences.
Artificial intelligence systems increasingly process video to assist humans in complex real-world tasks, making first-person, egocentric video comprehension vital. Egocentric video captures continuous, close-up interactions with tools and environments, requiring models to track fine-grained actions and objects over time. However, evaluating whether multi-modal large language models truly reason over time remains challenging because existing benchmarks often allow models to succeed through commonsense language guessing or static, single-frame visual clues rather than genuine temporal reasoning.
The article aims to rigorously evaluate and expose the temporal understanding limitations of modern multi-modal large language models in egocentric environments. To achieve this, it introduces EgoTempo, a new open-ended video question-answering benchmark designed specifically to require full-video comprehension across diverse first-person scenarios.
The authors constructed EgoTempo by extracting video clips averaging 45 seconds across 40 real-world scenario categories from the Ego4D dataset. Using a semi-automated pipeline combining vision-language models and manual filtering, they generated 500 validated question-answer pairs divided across 10 temporal capabilities covering actions and objects. They evaluated 13 leading commercial and open-source models, including Gemini, GPT-4o, Claude, and open-source architectures, assessing accuracy across text-only inputs, single frames, and multi-frame inputs sampled at varying rates.
The key findings reveal significant limitations in current model capabilities. First, the article demonstrates that existing egocentric benchmarks are heavily solvable without video: leading models achieve 41% to 51% accuracy on existing benchmarks using just a single frame, and up to 58% using only text, whereas on EgoTempo single-frame accuracy drops to just 9.1%. Second, current state-of-the-art models perform poorly on temporal reasoning even when provided with the full video: the highest-performing model, GPT-4o, achieved only 44.4% accuracy with 64 frames, and Gemini reached 39.1% at one frame per second. Third, while increasing frame counts improves performance from 9.1% to 39.1% (a 4.3-fold increase), accuracy quickly plateaus around 40%, showing that simply feeding more video frames into expanded context windows is insufficient. Fourth, tasks involving sequential ordering and counting proved especially difficult, showing little to no improvement from additional visual frames. Finally, a human baseline study achieved 63.2% average accuracy, outperforming the best model by roughly 24 percentage points and confirming a substantial machine-human performance gap.
These results imply that enterprise and research stakeholders should not assume that large context windows or high single-image benchmark scores translate to reliable video comprehension in time-critical, first-person applications such as robotics, manufacturing, or healthcare assistance. Relying on current models for sequential action verification or object tracking carries substantial operational risk due to frequent misinterpretations of temporal order and repeated actions.
The article recommends that AI development move beyond merely increasing input frame counts or context windows. Technical roadmaps should prioritize architectures and training techniques that explicitly model temporal dependencies, sequence transitions, and dynamic interactions. In addition, practitioners evaluating video models should adopt open-ended temporal benchmarks rather than standard multiple-choice datasets that disguise temporal reasoning failures through commonsense guessing.
Regarding confidence and limitations, the findings are supported by consistent evaluations across 13 distinct models and cross-validated automated grading showing a 96% alignment with human judgment. However, the benchmark scope is limited to 500 curated questions from 40 scenarios with an average duration of 45 seconds. Additionally, lower human performance on certain sequence tasks indicates inherent subjectivity in action granularity, meaning stakeholders should treat current performance numbers as conservative indicators of egocentric video reasoning complexity.
- Paper: Ego4D: Around the World in 3,000 Hours of Egocentric Video, Kristen Grauman et al. (2021). Ego4D establishes the egocentric video resource and benchmark context that underpins later work evaluating temporal understanding in first-person footage.
No sufficiently relevant recommendations were found.
