VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Shijie ZhouAlexander VilesovXuehai HeZiyu WanShuwang ZhangAditya NagachandraDi ChangDongdong ChenXin Eric WangAchuta Kadambi

article2025ICCV55 citations

Establishes VLM4D as a benchmark for evaluating spatiotemporal reasoning in vision language models across motion and perspective shifts, exposing critical performance gaps compared to humans and demonstrating how 4D feature field reconstruction improves dynamic scene understanding.

Listen

Modern vision-language models demonstrate impressive performance in static visual recognition and text generation, but they struggle significantly with dynamic visual environments. Real-world applications such as robotics, autonomous navigation, and physical artificial intelligence require systems to track moving objects, understand spatial rotations, compensate for camera perspectives, and maintain temporal continuity across time. While humans intuitively perceive dynamic events in four dimensions—combining three-dimensional space with time—existing artificial intelligence models primarily rely on aggregating two-dimensional features across frames. This fundamental architectural difference prevents current models from accurately reasoning about dynamic physical scenes, posing reliability and safety concerns for embodied deployment.

The article introduces and evaluates VLM4D, a dedicated benchmark designed to measure the spatiotemporal reasoning capabilities of vision-language models. The primary objective is to evaluate how well leading commercial and open-source models understand translational motion, rotational changes, perspective shifts, and temporal continuity, while exploring methods to bridge the gap between machine and human performance.

To establish a rigorous evaluation, the authors constructed a benchmark of 1,000 video sequences paired with over 1,800 human-verified question-answer pairs. The dataset combines third-person perspectives from standard video datasets, first-person egocentric recordings, and trajectory-controlled synthetic videos. The evaluation framework tests four core dimensions: translational motion, rotational movement, spatiotemporal counting, and false-positive event detection. The authors evaluated 23 state-of-the-art vision-language models ranging from 2 to 72 billion parameters under both direct output and step-by-step chain-of-thought reasoning modes.

The benchmark revealed a substantial performance gap between human baselines and artificial intelligence models. While human evaluators achieved an overall accuracy of 98.8%, the highest-performing model, Google's Gemini-2.5-Pro, achieved only 62.0%, followed by OpenAI's GPT-4o at 57.5%. Leading open-source models performed slightly lower, with top models achieving between 50.7% and 54.1% accuracy. Step-by-step chain-of-thought prompting provided no systematic benefit over direct answers, often introducing irrelevant details or internal logic contradictions. Furthermore, an audit of over two million existing video-tuning captions revealed that less than 10% of detected spatial-temporal labels accurately reflected true motion dynamics, indicating that current training datasets are severely deficient.

These findings indicate that relying on current vision-language models for spatial-temporal tasks carries substantial operational and safety risks. High performance on standard video captioning does not translate into robust physical understanding. To address these deficiencies, the authors demonstrated two promising solutions: targeted fine-tuning on temporally rich, high-quality dynamic data improved model accuracy by over 10 to 15 percentage points, while lifting two-dimensional video features into structured four-dimensional Gaussian feature fields during inference consistently boosted accuracy over standard video processing.

Organizations developing or deploying visual artificial intelligence systems should avoid relying purely on scale or off-the-shelf vision-language models for dynamic physical tasks. Instead, teams should invest in high-quality, motion-accurate spatiotemporal training datasets and explore four-dimensional scene representations. Further development is needed to make four-dimensional feature reconstruction computationally efficient, as current methods require scene-by-scene optimization. Because evaluations were conducted under zero-shot conditions on short video clips, stakeholders should exercise caution when extrapolating these results to unconstrained, long-duration operational environments.

No sufficiently relevant recommendations were found.

Cover for VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Abstract

Vision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions. Humans effortlessly track and reason about object movements, rotations, and perspective shifts-abilities essential for robust dynamic real-world understanding yet notably lacking in current VLMs. In this paper, we introduce VLM4D, the first benchmark specifically designed to evaluate the spatiotemporal reasoning capabilities of VLMs. Our benchmark comprises diverse real-world and synthetic videos accompanied by carefully curated question-answer pairs emphasizing translational and rotational motions, perspective awareness, and motion continuity. Through comprehensive evaluations of state-of-the-art open and closed-source VLMs, we identify significant performance gaps compared to human baselines, highlighting fundamental deficiencies in existing models. Extensive analysis reveals that VLMs struggle particularly with integrating multiple visual cues and maintaining temporal coherence. We further explore promising directions, such as leveraging 4D feature field reconstruction and targeted spatiotemporal supervised fine-tuning, demonstrating their effectiveness in enhancing spatiotemporal comprehension. Our work aims to encourage deeper exploration into improving VLMs' spatial and temporal grounding, paving the way towards more capable and reliable visual intelligence for dynamic environments.

Citation

MLA
Zhou, S., et al. “VLM4D: Towards Spatiotemporal Awareness in Vision Language Models”. arXiv, 2025, http://arxiv.org/abs/2508.02095v2.
APA
Zhou, S., Vilesov, A., He, X., Wan, Z., Zhang, S., Nagachandra, A., Chang, D., Chen, D., Wang, X. E., & Kadambi, A. (2025). VLM4D: Towards Spatiotemporal Awareness in Vision Language Models. arXiv. http://arxiv.org/abs/2508.02095v2
Chicago
Zhou, S., A. Vilesov, X. He, et al. 2025. “VLM4D: Towards Spatiotemporal Awareness in Vision Language Models”. arXiv. http://arxiv.org/abs/2508.02095v2.
Harvard
Zhou, S. et al. (2025) “VLM4D: Towards Spatiotemporal Awareness in Vision Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2508.02095v2.
Vancouver
1. Zhou S, Vilesov A, He X, Wan Z, Zhang S, Nagachandra A, Chang D, Chen D, Wang XE, Kadambi A (2025) VLM4D: Towards Spatiotemporal Awareness in Vision Language Models. arXiv

BibTeX

@article{zhou2025vlm4d,
  title = {VLM4D: Towards Spatiotemporal Awareness in Vision Language Models},
  author = {Zhou, Shijie and Vilesov, Alexander and He, Xuehai and Wan, Ziyu and Zhang, Shuwang and Nagachandra, Aditya and Chang, Di and Chen, Dongdong and Wang, Xin Eric and Kadambi, Achuta},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2508.02095v2},
  eprint = {2508.02095}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/