keyword
HourVideo
HourVideo is a multimodal benchmark dataset designed to evaluate the video-language understanding capabilities of artificial intelligence models on long-duration video content. Based on extended egocentric video footage lasting up to two hours, the benchmark assesses how effectively multimodal systems can process, retain, and analyze visual information across broad temporal contexts. It incorporates thousands of multiple-choice questions spanning diverse evaluation tasks, including video summarization, visual tracking and perception, spatial and causal reasoning, and environmental navigation. By probing long-form visual comprehension and temporal reasoning, HourVideo establishes a standardized framework for measuring the performance of long-context multimodal models in comparison to human-level comprehension.
1 item

