TAPVid-3D: A Benchmark for Tracking Any Point in 3D
Skanda KoppulaIgnacio RoccoYi YangJoseph HeywardJoão CarreiraAndrew ZissermanGabriel BrostowCarl Doersch
Introduces the first large-scale benchmark of over 4,000 real-world videos alongside standardized evaluation metrics to assess how accurately computer vision models track physical 3D point trajectories and occlusions across diverse dynamic scenes.
Enabling artificial intelligence and robotic systems to interact safely with the physical world requires models to accurately understand 3D scene geometry and dynamic motion from standard video. While two-dimensional point tracking benchmarks are well established, evaluating long-range three-dimensional point tracking (TAP-3D) in the real world has been impeded by a lack of diverse, real-world metric ground-truth datasets. Prior 3D evaluations have largely relied on synthetic simulations, which introduce a substantial domain gap and fail to capture real-world complexities.
The article introduces TAPVid-3D, a large-scale real-world benchmark designed to formalize and evaluate the ability of computer vision models to track any physical point in three dimensions across extended video sequences.
To construct the benchmark, the authors aggregated and standardized over 4,500 real-world video clips from three distinct sources: egocentric indoor manipulation footage from Aria Digital Twins, autonomous vehicle navigation data from DriveTrack, and dynamic multi-view human motion capture from Panoptic Studio. The dataset provides ground-truth metric coordinates and occlusion flags across solid surfaces, cleaned through automated filtering and manual verification. Alongside the dataset, the authors established unified evaluation metrics that measure 3D positional accuracy, occlusion classification, and overall tracking quality while adjusting for depth scale ambiguities.
The evaluation reveals a severe capability gap in existing vision systems when moving from 2D to 3D motion tracking. While top baseline models achieved strong 2D tracking performance—attaining up to 59.1% on standard 2D tracking metrics and over 85% occlusion accuracy—their performance dropped steeply to between 5.1% and 9.3% on the headline 3D Average Jaccard metric. Combining top 2D tracking algorithms with leading monocular depth estimators or multi-view reconstruction tools proved brittle, frequently suffering from scale drift over time and spatial inconsistencies across different objects. Newly emerging dedicated 3D tracking models achieved similar low performance, plateauing at 9.0% 3D Average Jaccard.
These findings indicate that existing perception systems cannot yet reliably serve as physical world models for embodied intelligence. For organizations building autonomous vehicles, robotic manipulators, and video generation tools, relying on current monocular 2D-to-3D tracking pipelines introduces significant performance and safety risks because apparent 2D tracking success masks severe 3D trajectory errors.
Engineering and research teams should utilize TAPVid-3D to develop and validate native 3D tracking architectures rather than relying on piecemeal combinations of 2D trackers and depth estimators. Practitioners should treat monocular 3D point tracking as an active research area rather than a production-ready technology for safety-critical spatial reasoning.
The benchmark has minor limitations, including sensor noise inherited from source datasets, domain concentration in specific urban driving and laboratory settings, and a restriction to opaque, solid objects. Nonetheless, it offers high confidence as a standardized, rigorous evaluation tool to guide progress in 3D physical scene understanding.
- Paper: DUSt3R: Geometric 3D Vision Made Easy, Shuzhe Wang et al. (2023). Introduces DUSt3R for uncalibrated dense 3D reconstruction and multi-view point alignment, representing the core feed-forward 3D vision paradigms evaluated and contrasted in TAPVid-3D.
- Paper: HOTA: A Higher Order Metric for Evaluating Multi-object Tracking, Jonathon Luiten et al. (2020). Defines Higher Order Tracking Accuracy (HOTA) to disentangle detection and association errors, establishing the benchmarking principles adapted for 3D point tracking evaluation.
- Paper: Object scene flow for autonomous vehicles, Moritz Menze et al. (2015). Establishes foundational datasets and metrics for dynamic 3D scene flow estimation in real-world driving environments.
- Paper: Unsupervised Learning of Depth and Ego-Motion from Video, Tinghui Zhou et al. (2017). Provides the foundational framework for learning monocular depth and ego-motion from video sequences that underpins standard 2D-to-3D tracking baselines.
- Paper: Deep Learning for 3D Point Clouds: A Survey, Yulan Guo et al. (2019). Surveys deep learning methodologies for irregular 3D point data and tracking representations across geometric benchmarks.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). Proposes a visual geometry grounded transformer architecture capable of directly predicting multi-frame point tracks and dense 3D geometry in a unified feed-forward pass.
- Paper: WildDet3D: Scaling Promptable 3D Detection in the Wild, Weikai Huang et al. (2026). Applies scalable 3D geometric prediction directly in unconstrained wild environments across arbitrary visual categories.
