VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
Shijie ZhouAlexander VilesovXuehai HeZiyu WanShuwang ZhangAditya NagachandraDi ChangDongdong ChenXin Eric WangAchuta Kadambi
Establishes VLM4D as a benchmark for evaluating spatiotemporal reasoning in vision language models across motion and perspective shifts, exposing critical performance gaps compared to humans and demonstrating how 4D feature field reconstruction improves dynamic scene understanding.
Modern vision-language models demonstrate impressive performance in static visual recognition and text generation, but they struggle significantly with dynamic visual environments. Real-world applications such as robotics, autonomous navigation, and physical artificial intelligence require systems to track moving objects, understand spatial rotations, compensate for camera perspectives, and maintain temporal continuity across time. While humans intuitively perceive dynamic events in four dimensions—combining three-dimensional space with time—existing artificial intelligence models primarily rely on aggregating two-dimensional features across frames. This fundamental architectural difference prevents current models from accurately reasoning about dynamic physical scenes, posing reliability and safety concerns for embodied deployment.
The article introduces and evaluates VLM4D, a dedicated benchmark designed to measure the spatiotemporal reasoning capabilities of vision-language models. The primary objective is to evaluate how well leading commercial and open-source models understand translational motion, rotational changes, perspective shifts, and temporal continuity, while exploring methods to bridge the gap between machine and human performance.
To establish a rigorous evaluation, the authors constructed a benchmark of 1,000 video sequences paired with over 1,800 human-verified question-answer pairs. The dataset combines third-person perspectives from standard video datasets, first-person egocentric recordings, and trajectory-controlled synthetic videos. The evaluation framework tests four core dimensions: translational motion, rotational movement, spatiotemporal counting, and false-positive event detection. The authors evaluated 23 state-of-the-art vision-language models ranging from 2 to 72 billion parameters under both direct output and step-by-step chain-of-thought reasoning modes.
The benchmark revealed a substantial performance gap between human baselines and artificial intelligence models. While human evaluators achieved an overall accuracy of 98.8%, the highest-performing model, Google's Gemini-2.5-Pro, achieved only 62.0%, followed by OpenAI's GPT-4o at 57.5%. Leading open-source models performed slightly lower, with top models achieving between 50.7% and 54.1% accuracy. Step-by-step chain-of-thought prompting provided no systematic benefit over direct answers, often introducing irrelevant details or internal logic contradictions. Furthermore, an audit of over two million existing video-tuning captions revealed that less than 10% of detected spatial-temporal labels accurately reflected true motion dynamics, indicating that current training datasets are severely deficient.
These findings indicate that relying on current vision-language models for spatial-temporal tasks carries substantial operational and safety risks. High performance on standard video captioning does not translate into robust physical understanding. To address these deficiencies, the authors demonstrated two promising solutions: targeted fine-tuning on temporally rich, high-quality dynamic data improved model accuracy by over 10 to 15 percentage points, while lifting two-dimensional video features into structured four-dimensional Gaussian feature fields during inference consistently boosted accuracy over standard video processing.
Organizations developing or deploying visual artificial intelligence systems should avoid relying purely on scale or off-the-shelf vision-language models for dynamic physical tasks. Instead, teams should invest in high-quality, motion-accurate spatiotemporal training datasets and explore four-dimensional scene representations. Further development is needed to make four-dimensional feature reconstruction computationally efficient, as current methods require scene-by-scene optimization. Because evaluations were conducted under zero-shot conditions on short video clips, stakeholders should exercise caution when extrapolating these results to unconstrained, long-duration operational environments.
- Paper: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark, Kunchang Li et al. (2023). MVBench established broad temporal video-understanding tasks for multimodal models, providing useful evaluation context for VLM4D’s more focused tests of motion and temporal coherence.
- Paper: ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos, Jr-Jen Chen et al. (2024). ReXTime shows how video benchmarks can isolate reasoning across temporally separated events, helping frame VLM4D’s evaluation of motion continuity and temporal coherence.
No sufficiently relevant recommendations were found.
