TAPVid-3D: A Benchmark for Tracking Any Point in 3D

Skanda KoppulaIgnacio RoccoYi YangJoseph HeywardJoão CarreiraAndrew ZissermanGabriel BrostowCarl Doersch

article2024NeurIPS53 citations

Introduces the first large-scale benchmark of over 4,000 real-world videos alongside standardized evaluation metrics to assess how accurately computer vision models track physical 3D point trajectories and occlusions across diverse dynamic scenes.

Listen

Enabling artificial intelligence and robotic systems to interact safely with the physical world requires models to accurately understand 3D scene geometry and dynamic motion from standard video. While two-dimensional point tracking benchmarks are well established, evaluating long-range three-dimensional point tracking (TAP-3D) in the real world has been impeded by a lack of diverse, real-world metric ground-truth datasets. Prior 3D evaluations have largely relied on synthetic simulations, which introduce a substantial domain gap and fail to capture real-world complexities.

The article introduces TAPVid-3D, a large-scale real-world benchmark designed to formalize and evaluate the ability of computer vision models to track any physical point in three dimensions across extended video sequences.

To construct the benchmark, the authors aggregated and standardized over 4,500 real-world video clips from three distinct sources: egocentric indoor manipulation footage from Aria Digital Twins, autonomous vehicle navigation data from DriveTrack, and dynamic multi-view human motion capture from Panoptic Studio. The dataset provides ground-truth metric coordinates and occlusion flags across solid surfaces, cleaned through automated filtering and manual verification. Alongside the dataset, the authors established unified evaluation metrics that measure 3D positional accuracy, occlusion classification, and overall tracking quality while adjusting for depth scale ambiguities.

The evaluation reveals a severe capability gap in existing vision systems when moving from 2D to 3D motion tracking. While top baseline models achieved strong 2D tracking performance—attaining up to 59.1% on standard 2D tracking metrics and over 85% occlusion accuracy—their performance dropped steeply to between 5.1% and 9.3% on the headline 3D Average Jaccard metric. Combining top 2D tracking algorithms with leading monocular depth estimators or multi-view reconstruction tools proved brittle, frequently suffering from scale drift over time and spatial inconsistencies across different objects. Newly emerging dedicated 3D tracking models achieved similar low performance, plateauing at 9.0% 3D Average Jaccard.

These findings indicate that existing perception systems cannot yet reliably serve as physical world models for embodied intelligence. For organizations building autonomous vehicles, robotic manipulators, and video generation tools, relying on current monocular 2D-to-3D tracking pipelines introduces significant performance and safety risks because apparent 2D tracking success masks severe 3D trajectory errors.

Engineering and research teams should utilize TAPVid-3D to develop and validate native 3D tracking architectures rather than relying on piecemeal combinations of 2D trackers and depth estimators. Practitioners should treat monocular 3D point tracking as an active research area rather than a production-ready technology for safety-critical spatial reasoning.

The benchmark has minor limitations, including sensor noise inherited from source datasets, domain concentration in specific urban driving and laboratory settings, and a restriction to opaque, solid objects. Nonetheless, it offers high confidence as a standardized, rigorous evaluation tool to guide progress in 3D physical scene understanding.

Cover for TAPVid-3D: A Benchmark for Tracking Any Point in 3D

Abstract

We introduce a new benchmark, TAPVid-3D, for evaluating the task of long-range Tracking Any Point in 3D (TAP-3D). While point tracking in two dimensions (TAP-2D) has many benchmarks measuring performance on real-world videos, such as TAPVid-DAVIS, three-dimensional point tracking has none. To this end, leveraging existing footage, we build a new benchmark for 3D point tracking featuring 4,000+ real-world videos, composed of three different data sources spanning a variety of object types, motion patterns, and indoor and outdoor environments. To measure performance on the TAP-3D task, we formulate a collection of metrics that extend the Jaccard-based metric used in TAP-2D to handle the complexities of ambiguous depth scales across models, occlusions, and multi-track spatio-temporal smoothness. We manually verify a large sample of trajectories to ensure correct video annotations, and assess the current state of the TAP-3D task by constructing competitive baselines using existing tracking models. We anticipate this benchmark will serve as a guidepost to improve our ability to understand precise 3D motion and surface deformation from monocular video.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 TAPVid-3D
  • 3.1 Aria Digital Twins
  • 3.2 DriveTrack
  • 3.3 Panoptic Studio
  • 3.4 Data Cleanup and Validation
  • 3.5 Metrics
  • 4 Baselines on TAPVid-3D
  • 5 Conclusion
  • References
  • NeurIPS Paper Checklist

Knowls

  1. Knowl 1 — Tracking Any Point in 3D (TAP-3D) Task Formulation

    definition

    The Tracking Any Point in 3D (TAP-3D) task formalizes long-range motion and 3D geometry estimation of arbitrary physical points from a monocular RGB video sequence. Given an input video V={It}t=1TV = \{I_t\}_{t=1}^T consisting of TT frames of spatial dimensions W×HW \times H, camera intrinsic calibration matrix K∈R3×3K \in \mathbb{R}^{3 \times 3}, and a query point q=(xq,yq,tq)q = (x_q, y_q, t_q) specifying a 2D pixel location (xq,yq)(x_q, y_q) observed at query frame tq∈{1,…,T}t_q \in \{1, \dots, T\}, a model must predict for every frame t∈{1,…,T}t \in \{1, \dots, T\}:

    1. The metric 3D point position P^t=(x^t,y^t,z^t)∈R3\hat{P}_t = (\hat{x}_t, \hat{y}_t, \hat{z}_t) \in \mathbb{R}^3 of the physical query point expressed in the camera coordinate frame (Qcam)t(Q_{\text{cam}})_t.
    2. A binary visibility indicator v^t∈{0,1}\hat{v}_t \in \{0, 1\}, where v^t=1\hat{v}_t = 1 denotes that the physical point is visible within frame ItI_t, and v^t=0\hat{v}_t = 0 denotes that the point is occluded or outside the image boundary.

    By construction, the point is visible at its query timestamp (v^tq=1\hat{v}_{t_q} = 1). Unlike instantaneous scene flow or 2D tracking, TAP-3D tracks arbitrary points over extended temporal sequences (spanning tens to hundreds of frames) across rigid, articulated, and non-rigid surface deformations without relying on category-specific 3D mesh models or priors.

  2. Knowl 2 — TAP-3D Evaluation Metrics: APD_3D, OA, and AJ_3D

    definition

    To evaluate metric 3D point tracking accuracy, point visibility, and combined tracking performance, TAP-3D adapts 2D TAP metrics using a depth-adaptive distance threshold.

    Let Pti∈R3P^i_t \in \mathbb{R}^3 and P^ti∈R3\hat{P}^i_t \in \mathbb{R}^3 be the ground-truth and predicted metric 3D camera-frame positions for trajectory ii at frame tt, and let vti∈{0,1}v^i_t \in \{0, 1\} and v^ti∈{0,1}\hat{v}^i_t \in \{0, 1\} be the ground-truth and predicted visibility flags.

    Depth-Adaptive Threshold: Because tracking errors closer to the camera represent larger visual and physical displacements relative to distant points, the 3D error threshold δ3D(Pti)\delta_{3D}(P^i_t) is defined by back-projecting 2D pixel thresholds δ2D∈{1,2,4,8,16}\delta_{2D} \in \{1, 2, 4, 8, 16\} pixels using focal length ff and ground-truth depth Z(Pti)Z(P^i_t): δ3D(Pti)=Z(Pti)⋅δ2Df\delta_{3D}(P^i_t) = \frac{Z(P^i_t) \cdot \delta_{2D}}{f}

    Average Position Error within Delta (APD3D\text{APD}_{3D}): Measures the percentage of visible points whose predicted 3D position falls within the distance threshold δ3D(Pti)\delta_{3D}(P^i_t): APD3D≡1V∑i,tvti⋅I(∥P^ti−Pti∥2<δ3D(Pti))\text{APD}_{3D} \equiv \frac{1}{V} \sum_{i, t} v^i_t \cdot \mathbb{I}\left(\| \hat{P}^i_t - P^i_t \|_2 < \delta_{3D}(P^i_t)\right) where I(⋅)\mathbb{I}(\cdot) is the indicator function, V=∑i,tvtiV = \sum_{i, t} v^i_t is the total count of visible points across all trajectories and frames, and the final value is averaged across the 5 threshold values δ2D\delta_{2D}.

    Occlusion Accuracy (OA\text{OA}): Measures the classification accuracy of point visibility across all frames: OA≡1Ntotal∑i,tI(v^ti=vti)\text{OA} \equiv \frac{1}{N_{\text{total}}} \sum_{i, t} \mathbb{I}\left(\hat{v}^i_t = v^i_t\right) where NtotalN_{\text{total}} is the total number of evaluated trajectory points.

    3D Average Jaccard (AJ3D\text{AJ}_{3D}): The overall headline tracking metric balances position precision and occlusion classification: AJ3D≡∑i,tvti v^ti αti∑i,tvti+∑i,t(1−vti) v^ti+∑i,tvti v^ti (1−αti)\text{AJ}_{3D} \equiv \frac{\sum_{i, t} v^i_t \, \hat{v}^i_t \, \alpha^i_t}{\sum_{i, t} v^i_t + \sum_{i, t} (1 - v^i_t)\,\hat{v}^i_t + \sum_{i, t} v^i_t \, \hat{v}^i_t \, (1 - \alpha^i_t)} where αti=I(∥P^ti−Pti∥2<δ3D(Pti))\alpha^i_t = \mathbb{I}\left(\| \hat{P}^i_t - P^i_t \|_2 < \delta_{3D}(P^i_t)\right) indicates whether the 3D position error satisfies the distance threshold.

  3. Knowl 3 — Depth Scale Rescaling Protocols for Monocular TAP-3D

    model/method

    Because monocular depth estimation and tracking models often suffer from scale ambiguity, predictions are aligned with ground truth prior to computing metric tracking performance using one of two rescaling protocols:

    1. Global Median Rescaling: All predicted 3D points P^ti\hat{P}^i_t across all trajectories ii and frames tt in a video are scaled by a single global scalar factor sglobals_{\text{global}}, defined as the median ratio of ground truth to predicted 3D norms: sglobal=mediani,t(∥Pti∥2∥P^ti∥2)s_{\text{global}} = \text{median}_{i, t} \left( \frac{\| P^i_t \|_2}{\| \hat{P}^i_t \|_2} \right) P^ti,rescaled=sglobal⋅P^ti\hat{P}^{i, \text{rescaled}}_t = s_{\text{global}} \cdot \hat{P}^i_t This protocol evaluates the model's ability to maintain a globally consistent 3D scale across both time and space.

    2. Per-Trajectory Rescaling: For models that estimate relative depth changes for individual points accurately but lack consistent global scale across separate points (e.g., models trained exclusively on synthetic simulation data), each trajectory ii is scaled independently using the scale at its query timestamp tqt_q: straji=∥Ptqi∥2∥P^tqi∥2s^i_{\text{traj}} = \frac{\| P^i_{t_q} \|_2}{\| \hat{P}^i_{t_q} \|_2} P^ti,rescaled=straji⋅P^ti∀t\hat{P}^{i, \text{rescaled}}_t = s^i_{\text{traj}} \cdot \hat{P}^i_t \quad \forall t

  4. Knowl 4 — TAPVid-3D Benchmark Dataset Composition and Statistics

    experimental setup

    TAPVid-3D is a monocular real-world benchmark containing 4,569 video clips spanning 2,828 videos and 255 unique scenes, derived from three distinct data domains:

    • Aria Digital Twins (ADT): Egocentric video captured with Aria glasses in indoor studio environments with paired digital replicas, representing first-person manipulation and camera-hand interactions.
    • DriveTrack: Forward-facing driving videos from the Waymo Open Dataset across 6 US cities (San Francisco, Phoenix, Mountain View, Los Angeles, Detroit, Seattle), representing outdoor rigid vehicle motion.
    • Panoptic Studio: Massively multi-view stationary camera dome footage depicting humans performing dynamic, complex non-rigid actions.

    The benchmark includes two standardized splits:

    • minival: 150 clips (50 clips per data source) designated for lightweight online validation during training.
    • full_test: The complete test benchmark comprising all 4,569 clips (excluding the minival clips).
    Dataset split #clips (minival) #trajs/clip #frames/clip #videos #scenes Resolution FPS
    Aria Digital Twins 1956 (50) 1024 300 215 2 512×512512 \times 512 30
    DriveTrack 2457 (50) 256 25−30025 - 300 2457 252 1920×12801920 \times 1280 10
    Panoptic Studio 156 (50) 50 150 156 1 640×360640 \times 360 30
    TAPVid-3D (Total) 4569 (150) 50−102450 - 1024 25−30025 - 300 2828 255 Multiple 10 / 30
  5. Knowl 5 — Aria Digital Twins 3D Trajectory Annotation Pipeline

    model/method

    In the Aria Digital Twins split of TAPVid-3D, metric 3D point tracks and visibility annotations are constructed from digital studio replicas.

    Given a video V={It}t=1TV = \{I_t\}_{t=1}^T, depth maps {Dt}⊂RW×H\{D_t\} \subset \mathbb{R}^{W \times H}, object instance segmentation masks {St}⊂ZW×H\{S_t\} \subset \mathbb{Z}^{W \times H}, camera intrinsic matrix K∈R3×3K \in \mathbb{R}^{3 \times 3}, world-to-camera poses {(Pcamw)t}⊂R3×4\{(P^{\text{w}}_{\text{cam}})_t\} \subset \mathbb{R}^{3 \times 4}, and a query point q=(xq,yq,tq)q = (x_q, y_q, t_q):

    1. The 3D position in the camera coordinate frame at query timestamp tqt_q, (Qcam)tq(Q_{\text{cam}})_{t_q}, is computed by unprojecting using depth Dtq(xq,yq)D_{t_q}(x_q, y_q): (Qcam)tq=K−1(xqyq1)⋅Dtq(xq,yq)(Q_{\text{cam}})_{t_q} = K^{-1} \begin{pmatrix} x_q \\ y_q \\ 1 \end{pmatrix} \cdot D_{t_q}(x_q, y_q)

    2. The object ID qid=Stq(xq,yq)q_{\text{id}} = S_{t_q}(x_q, y_q) is used to look up the 3D object pose (Pobjw)tq(P^{\text{w}}_{\text{obj}})_{t_q}, transforming the point into object coordinates QobjQ_{\text{obj}}: Qobj=(Pobjw)tq(Pwcam)tq(Qcam)tqQ_{\text{obj}} = (P^{\text{w}}_{\text{obj}})_{t_q} (P^{\text{cam}}_{\text{w}})_{t_q} (Q_{\text{cam}})_{t_q}

    3. For any frame tt, the 3D camera-frame position is retrieved by: (Qcam)t=(Pcamw)t(Pwobj)tQobj(Q_{\text{cam}})_t = (P^{\text{w}}_{\text{cam}})_t (P^{\text{obj}}_{\text{w}})_t Q_{\text{obj}}

    4. Visibility vt∈{0,1}v_t \in \{0, 1\} is computed by projecting (Qcam)t(Q_{\text{cam}})_t to image coordinates (u,v)=ΠK((Qcam)t)(u, v) = \Pi_K((Q_{\text{cam}})_t), testing whether the camera-frame depth Z((Qcam)t)Z((Q_{\text{cam}})_t) agrees with Dt(u,v)D_t(u, v) within threshold δ\delta, and checking whether the point is occluded by operator hands using a semantic hand segmentation mask Ht(u,v)∈{0,1}H_t(u, v) \in \{0, 1\}: vt=I(∣Z((Qcam)t)−Dt(u,v)∣<δ)⋅(1−Ht(u,v))v_t = \mathbb{I}\left(|Z((Q_{\text{cam}})_t) - D_t(u, v)| < \delta\right) \cdot \left(1 - H_t(u, v)\right)

  6. Knowl 6 — DriveTrack and Panoptic Studio 3D Trajectory Annotation Pipelines

    model/method

    Ground truth metric 3D point trajectories for DriveTrack and Panoptic Studio are extracted using sensor-specific pipelines:

    DriveTrack Split:

    1. At sample time tst_s, points (Qcam)ts(Q_{\text{cam}})_{t_s} from Waymo LIDAR point clouds corresponding to a single tracked vehicle are extracted via annotated 3D bounding boxes.
    2. Assuming vehicle rigidity, points are mapped into the object coordinate frame QobjQ_{\text{obj}} and tracked across all frames tt via the vehicle's bounding box poses and camera extrinsic poses.
    3. Dense depth maps DtD_t are obtained by interpolating sparse LIDAR returns. Point visibility vtv_t is assigned 0 if the query point's distance from the camera center exceeds Dt(u,v)D_t(u, v) by more than a 5% relative threshold margin.
    4. Query timestamp tqt_q is sampled uniformly from visible frames, and query pixel coordinates are computed via (xq,yq)=ΠK((Qcam)tq)(x_q, y_q) = \Pi_K((Q_{\text{cam}})_{t_q}).

    Panoptic Studio Split:

    1. Dynamic 3D Gaussian Splatting models fitted across dome cameras represent non-rigid actor motions via a set of 3D Gaussians {(μi,Σi)t0}i=1N\{(\mu_i, \Sigma_i)_{t_0}\}_{i=1}^N displaced over time.
    2. Given 2D query point q=(xq,yq,tq)q = (x_q, y_q, t_q), (Qcam)tq(Q_{\text{cam}})_{t_q} is unprojected and matched to the nearest Gaussian center index i∗=arg⁡min⁡i∥(Qcam)tq−(Pcamw)tq(μi)tq∥2i^* = \arg\min_i \| (Q_{\text{cam}})_{t_q} - (P^{\text{w}}_{\text{cam}})_{t_q} (\mu_i)_{t_q} \|_2.
    3. The 2D query is adjusted to (xq,yq)=ΠK((Pcamw)tq(μi∗)tq)(x_q, y_q) = \Pi_K((P^{\text{w}}_{\text{cam}})_{t_q} (\mu_{i^*})_{t_q}), and its 3D trajectory across all frames is defined by the motion of that Gaussian center: (Qcam)t=(Pcamw)t(μi∗)t(Q_{\text{cam}})_t = (P^{\text{w}}_{\text{cam}})_t (\mu_{i^*})_t.
    4. Visibility vtv_t is computed by verifying agreement between point depth and rendered depth maps DtD_t. Tracking is restricted exclusively to foreground dynamic actors.

    Automated Cleanup & Validation: Trajectories exceeding frame-wise instance segmentation boundaries (via Segment Anything) when unoccluded are filtered out (removing 2–3% of initial tracks in DriveTrack). Trajectories exhibiting visibility flickering (visibility state changing more than 10% of the total frame count) are removed.

  7. Knowl 7 — Benchmark Baseline Performance and the 2D vs. 3D Tracking Gap

    empirical result

    Evaluating baseline models constructed by pairing 2D point trackers (TAPIR, CoTracker, BootsTAPIR) with monocular depth estimators (ZoeDepth, Depth Anything V2) and Structure-from-Motion (COLMAP) demonstrates a severe performance drop when transitioning from 2D tracking to 3D metric point tracking on TAPVid-3D.

    Aria DriveTrack PStudio Average
    Baseline Model 3D-AJ APD OA 3D-AJ APD OA 3D-AJ APD OA 3D-AJ APD OA
    Static Baseline 4.9 10.2 55.4 3.9 6.5 80.8 5.9 11.5 75.8 4.9 9.4 70.7
    TAPIR + COLMAP 7.1 11.9 72.6 8.9 14.7 80.4 6.1 10.7 75.2 7.4 12.4 76.1
    CoTracker + COLMAP 8.0 12.3 78.6 11.7 19.1 81.7 8.1 13.5 77.2 9.3 15.0 79.1
    BootsTAPIR + COLMAP 9.1 14.5 78.6 11.8 18.6 83.8 6.9 11.6 81.8 9.3 14.9 81.4
    TAPIR + ZoeDepth 9.0 14.3 79.7 5.2 8.8 81.6 10.7 18.2 78.7 8.3 13.8 80.0
    CoTracker + ZoeDepth 10.0 15.9 87.8 5.0 9.1 82.6 11.2 19.4 80.0 8.7 14.8 83.4
    BootsTAPIR + ZoeDepth 9.9 16.3 86.5 5.4 9.2 85.3 11.3 19.0 82.7 8.8 14.8 84.8
    BootsTAPIR + DepthAnythingV2 3.7 6.9 86.5 7.4 12.4 85.4 5.6 10.3 82.7 5.5 9.9 84.9
    CoTracker + DepthAnythingV2 3.6 6.6 87.8 7.1 12.6 82.6 5.6 10.5 80.0 5.4 9.9 83.5
    TAPIR + DepthAnythingV2 3.3 6.1 79.6 6.9 11.5 81.7 5.3 9.8 78.7 5.1 9.1 80.0
    TAPIR-3D 2.5 4.8 86.0 3.2 5.9 83.3 3.6 7.0 78.9 3.1 5.9 82.8
    SpatialTracker 9.9 16.1 89.0 6.2 11.1 83.7 10.9 19.2 78.6 9.0 15.5 83.7
    BootsTAPIR + COLMAP* 7.3 11.5 76.3 9.3 15.1 83.5 6.2 10.6 78.7 7.6 12.4 79.5
    BootsTAPIR + ZoeDepth* 8.6 14.5 86.9 5.1 8.7 83.5 10.2 17.7 82.0 8.0 13.6 84.1
    SpatialTracker* 9.2 15.1 89.9 5.8 10.2 82.0 9.8 17.7 78.4 8.3 14.3 83.4

    *Rows marked with an asterisk indicate evaluation on the 150-clip minival split.

    When evaluated purely on 2D tracking by projecting ground truth 3D tracks onto the image plane, these models achieve competitive scores across TAPVid-3D:

    • TAPIR: 2D-AJ=53.2%\text{2D-AJ} = 53.2\%, APD=67.4%\text{APD} = 67.4\%, OA=80.5%\text{OA} = 80.5\%
    • CoTracker: 2D-AJ=57.2%\text{2D-AJ} = 57.2\%, APD=74.2%\text{APD} = 74.2\%, OA=84.5%\text{OA} = 84.5\%
    • BootsTAPIR: 2D-AJ=59.1%\text{2D-AJ} = 59.1\%, APD=74.7%\text{APD} = 74.7\%, OA=85.6%\text{OA} = 85.6\%

    In contrast, their average 3D tracking performance (3D-AJ\text{3D-AJ}) remains below 10.0%10.0\%. The key failure modes in 3D tracking include frame-to-frame depth estimation noise, temporal depth scale drift over long trajectories, and spatial depth inconsistency across distinct points preventing global scale alignment.

  8. Knowl 8 — Limitations of the TAPVid-3D Benchmark

    limitation

    The TAPVid-3D benchmark has several identified limitations:

    1. Domain Scope: The benchmark is constructed from three specific domains (indoor studio replicas for Aria Digital Twins, autonomous driving in 6 US cities for DriveTrack, and an indoor capture dome for Panoptic Studio), which do not span all real-world dynamics, viewpoints, and environments.
    2. Material Constraints: Tracking evaluation is restricted exclusively to solid, opaque surfaces, inheriting limitations from monocular depth and standard TAP formulations. Transparent, reflective, refractive, or volumetric visual phenomena (e.g., water, smoke, glass) are not evaluated.
    3. Annotation Noise: Despite automated filtering via instance segmentation and visibility flickering thresholds, residual noise can occur from slight misalignments between physical videos and synthetic CAD replicas in Aria Digital Twins, LIDAR sparsity and occlusion boundaries in DriveTrack, and dynamic Gaussian splat drift in Panoptic Studio.
    4. Geographic and Demographic Biases: Videos originate from lab researchers or driving in specific urban and suburban US cities, inheriting demographic and geographic biases present in the underlying data sources.

Coverage note — Omitted background and related work surveys (such as general summaries of NRSfM, monocular depth estimation literature, and 3D pose tracking) and standard checklist questionnaire responses, as these do not constitute novel contributions of the paper.

References

  1. 1.Introducing Project Aria Glasses, from Meta. URL https://www.projectaria.com/glasses/.
  2. 2.Ijaz Akhter, Yaser Sheikh, Sohaib Khan, and Takeo Kanade. Nonrigid structure from motion in trajectory space. In D. Koller, D. Schuurmans, Y. Bengio, and L. Bottou, editors, NeurIPS, 2008.
  3. 3.Arjun Balasingam, Joseph Chandler, Chenning Li, Zhoutong Zhang, and Hari Balakrishnan. Drivetrack: A benchmark for long-range point tracking in real-world videos. arXiv preprint arXiv:2312.09523, 2023.
  4. 4.Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias Müller. Zoedepth: Zero-shot transfer by combining relative and metric depth. arXiv preprint arXiv:2302.12288, 2023.
  5. 5.Benjamin Biggs, Thomas Roddick, Andrew Fitzgibbon, and Roberto Cipolla. Creatures great and SMAL: Recovering the shape and motion of animals from video. In Proc. ACCV, 2019.
  6. 6.Christoph Bregler, Aaron Hertzmann, and Henning Biermann. Recovering non-rigid 3d shape from image streams. In Proc. CVPR, 2000.
  7. 7.Daniel J Butler, Jonas Wulff, Garrett B Stanley, and Michael J Black. A naturalistic open source movie for optical flow evaluation. In Computer Vision–ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part VI 12, pages 611–625. Springer, 2012.
  8. 8.Vincent Casser, Soeren Pirk, Reza Mahjourian, and Anelia Angelova. Depth prediction without the sensors: Leveraging structure for unsupervised learning from monocular videos. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 8001–8008, 2019.
  9. 9.Haiyang Chao, Yu Gu, and Marcello Napolitano. A survey of optical flow techniques for robotics navigation applications. Journal of Intelligent & Robotic Systems, 73:361–372, 2014.
  10. 10.Carl Doersch, Ankush Gupta, Larisa Markeeva, Adrià Recasens, Lucas Smaira, Yusuf Aytar, João Carreira, Andrew Zisserman, and Yi Yang. TAP-Vid: A benchmark for tracking any point in a video. NeurIPS, 2022.
  11. 11.Carl Doersch, Yi Yang, Mel Vecerik, Dilara Gokay, Ankush Gupta, Yusuf Aytar, Joao Carreira, and Andrew Zisserman. TAPIR: Tracking any point with per-frame initialization and temporal refinement. arXiv preprint arXiv:2306.08637, 2023.
  12. 12.Carl Doersch, Yi Yang, Dilara Gokay, Pauline Luc, Skanda Koppula, Ankush Gupta, Joseph Heyward, Ross Goroshin, João Carreira, and Andrew Zisserman. Bootstap: Bootstrapped training for tracking-any-point. arXiv preprint arXiv:2402.00847, 2024.
  13. 13.Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, and Jitendra Malik. Humans in 4d: Reconstructing and tracking humans with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14783–14794, 2023.
  14. 14.Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, et al. Kubric: A scalable dataset generator. In Proc. CVPR, 2022.
  15. 15.Yang Hai, Rui Song, Jiaojiao Li, Mathieu Salzmann, and Yinlin Hu. Rigidity-aware detection for 6d object pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8927–8936, 2023.
  16. 16.Albert Haque, Boya Peng, Zelun Luo, Alexandre Alahi, Serena Yeung, and Li Fei-Fei. Towards viewpoint invariant 3d human pose estimation. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14, pages 160–177. Springer, 2016.
  17. 17.Adam W Harley, Zhaoyuan Fang, and Katerina Fragkiadaki. Particle video revisited: Tracking through occlusions using point trajectories. In Proc. ECCV, 2022.
  18. 18.Chong Huang and Kazuhito Koishida. Improved active speaker detection based on optical flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 950–951, 2020.
  19. 19.H. Jhuang, J. Gall, S. Zuffi, C. Schmid, and M. J. Black. Towards understanding action recognition. In International Conf. on Computer Vision (ICCV), pages 3192–3199, 2013.
  20. 20.Sam Johnson and Mark Everingham. Clustered pose and nonlinear appearance models for human pose estimation. In bmvc, volume 2, page 5. Aberystwyth, UK, 2010.
  21. 21.Hanbyul Joo, Tomas Simon, Xulong Li, Hao Liu, Lei Tan, Lin Gui, Sean Banerjee, and Timothy Godisart. Panoptic studio: A massively multiview system for social interaction capture. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(1), 2019.
  22. 22.Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. CoTracker: It is better to track together. arXiv preprint arXiv:2307.07635, 2023.
  23. 23.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 2023.
  24. 24.Ishan Khatri, Kyle Vedder, Neehar Peri, Deva Ramanan, and James Hays. I can’t believe it’s not scene flow! arXiv preprint arXiv:2403.04739, 2024.
  25. 25.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4015–4026, 2023.
  26. 26.Johannes Kopf, Xuejian Rong, and Jia-Bin Huang. Robust consistent video depth estimation. In Proc. CVPR, 2021.
  27. 27.Jiefeng Li, Can Wang, Hao Zhu, Yihuan Mao, Hao-Shu Fang, and Cewu Lu. Crowdpose: Efficient crowded scenes pose estimation and a new benchmark. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10863–10872, 2019.
  28. 28.Zhengqi Li and Noah Snavely. Megadepth: Learning singleview depth prediction from internet photos. ieee. In Proc. CVPR, 2018.
  29. 29.Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view synthesis of dynamic scenes. In Proc. CVPR, 2021.
  30. 30.Zhengqi Li, Qianqian Wang, Forrester Cole, Richard Tucker, and Noah Snavely. Dynibar: Neural dynamic image-based rendering. In Proc. CVPR, 2023.
  31. 31.Yiyi Liao, Jun Xie, and Andreas Geiger. Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d. IEEE PAMI, 45(3), 2022.
  32. 32.Joseph J Lim, Hamed Pirsiavash, and Antonio Torralba. Parsing ikea objects: Fine pose estimation. In Proceedings of the IEEE international conference on computer vision, pages 2992–2999, 2013.
  33. 33.Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. Hoi4d: A 4d egocentric dataset for category-level human-object interaction. In Proc. CVPR, 2022.
  34. 34.Manolis Lourakis and Xenophon Zabulis. Model-based pose estimation for rigid objects. In International conference on computer vision systems, pages 83–92. Springer, 2013.
  35. 35.Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. In 3DV, 2024.
  36. 36.Xuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen, and Johannes Kopf. Consistent video depth estimation. ACM Transactions on Graphics, 2020.
  37. 37.Wei-Chiu Ma, Shenlong Wang, Rui Hu, Yuwen Xiong, and Raquel Urtasun. Deep rigid instance scene flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3614–3622, 2019.
  38. 38.Alexander Mathis, Thomas Biasi, Steffen Schneider, Mert Yuksekgonul, Byron Rogers, Matthias Bethge, and Mackenzie W Mathis. Pretraining boosts out-of-domain robustness for pose estimation. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 1859–1868, 2021.
  39. 39.Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4040–4048, 2016.
  40. 40.Moritz Menze and Andreas Geiger. Object scene flow for autonomous vehicles. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3061–3070, 2015.
  41. 41.Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters, Thomas Whelan, Chen Kong, Omkar Parkhi, Richard Newcombe, and Yuheng Carl Ren. Aria digital twin: A new benchmark dataset for egocentric 3d machine perception. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 20133–20143, 2023.
  42. 42.Viorica Patraucean, Lucas Smaira, Ankush Gupta, Adria Recasens, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Mateusz Malinowski, Yi Yang, Carl Doersch, et al. Perception test: A diagnostic benchmark for multimodal video models. 2024.
  43. 43.René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE PAMI, 2020.
  44. 44.Johannes L Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In Proc. CVPR, 2016.
  45. 45.Johannes L Schonberger, Hans Hardmeier, Torsten Sattler, and Marc Pollefeys. Comparative evaluation of hand-crafted and learned local features. In Proc. CVPR, pages 1482–1491, 2017.
  46. 46.Johannes Lutz Schönberger and Jan-Michael Frahm. Structure-from-motion revisited. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  47. 47.Johannes Lutz Schönberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. Pixelwise view selection for unstructured multi-view stereo. In European Conference on Computer Vision (ECCV), 2016.
  48. 48.Cristian Sminchisescu. 3d human motion analysis in monocular video: techniques and challenges. Human Motion: Understanding, Modelling, Capture, and Animation, pages 185–211, 2008.
  49. 49.Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 567–576, 2015.
  50. 50.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2446–2454, 2020.
  51. 51.Ramana Sundararaman, Cedric De Almeida Braga, Eric Marchand, and Julien Pettre. Tracking pedestrian heads in dense crowd. In Proc. CVPR, 2021.
  52. 52.Jonathan Tompson, Murphy Stein, Yann Lecun, and Ken Perlin. Real-time continuous pose recovery of human hands using convolutional networks. ACM Transactions on Graphics, 33, August 2014.
  53. 53.Mel Vecerik, Carl Doersch, Yi Yang, Todor Davchev, Yusuf Aytar, Guangyao Zhou, Raia Hadsell, Lourdes Agapito, and Jon Scholz. RoboTAP: Tracking arbitrary points for few-shot visual imitation. In Proc. Intl. Conf. on Robotics and Automation, 2024.
  54. 54.Sundar Vedula, Simon Baker, Peter Rander, Robert Collins, and Takeo Kanade. Three-dimensional scene flow. In Proc. ICCV, 1999.
  55. 55.Timo Von Marcard, Roberto Henschel, Michael J Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering accurate 3d human pose in the wild using imus and a moving camera. In Proceedings of the European conference on computer vision (ECCV), pages 601–617, 2018.
  56. 56.Bo Wang, Jian Li, Yang Yu, Li Liu, Zhenping Sun, and Dewen Hu. SceneTracker: Long-term scene flow estimation network. arXiv preprint arXiv:2403.19924, 2024.
  57. 57.Zhouxia Wang, Ziyang Yuan, Xintao Wang, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. Motionctrl: A unified and flexible motion controller for video generation. arXiv preprint arXiv:2312.03641, 2023.
  58. 58.Chuan Wen, Xingyu Lin, John So, Kai Chen, Qi Dou, Yang Gao, and Pieter Abbeel. Any-point trajectory modeling for policy learning. arXiv preprint arXiv:2401.00025, 2023.
  59. 59.Benjamin Wilson, William Qi, Tanmay Agarwal, John Lambert, Jagjeet Singh, Siddhesh Khandelwal, Bowen Pan, Ratnesh Kumar, Andrew Hartnett, Jhony Kaesemodel Pontes, et al. Argoverse 2: Next generation datasets for self-driving perception and forecasting. arXiv preprint arXiv:2301.00493, 2023.
  60. 60.Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, and Xiaowei Zhou. Spatialtracker: Tracking any 2d pixels in 3d space. arXiv preprint arXiv:2404.04319, 2024.
  61. 61.Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. arXiv preprint arXiv:2401.10891, 2024.
  62. 62.Zeyu Yang, Hongye Yang, Zijie Pan, Xiatian Zhu, and Li Zhang. Real-time photorealistic dynamic scene representation and rendering with 4d gaussian splatting. arXiv preprint arXiv:2310.10642, 2023.
  63. 63.Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. Mvsnet: Depth inference for unstructured multi-view stereo. In Proc. ECCV, pages 767–783, 2018.
  64. 64.Emilie Yu, Kevin Blackburn-Matzen, Cuong Nguyen, Oliver Wang, Rubaiat Habib Kazi, and Adrien Bousseau. VideoDoodles: Hand-drawn animations on videos with scene-aware canvases. ACM Transactions on Graphics, 42(4):1–12, 2023.
  65. 65.Hang Yu, Yufei Xu, Jing Zhang, Wei Zhao, Ziyu Guan, and Dacheng Tao. Ap-10k: A benchmark for animal pose estimation in the wild. arXiv preprint arXiv:2108.12617, 2021.
  66. 66.Weiyu Zhang, Menglong Zhu, and Konstantinos G Derpanis. From actemes to action: A strongly-supervised representation for detailed action understanding. In Proceedings of the IEEE international conference on computer vision, pages 2248–2255, 2013.
  67. 67.Zhoutong Zhang, Forrester Cole, Richard Tucker, William T Freeman, and Tali Dekel. Consistent depth of moving objects in video. ACM Transactions on Graphics, 2021.
  68. 68.Yang Zheng, Adam W Harley, Bokui Shen, Gordon Wetzstein, and Leonidas J Guibas. PointOdyssey: A large-scale synthetic dataset for long-term point tracking. In Proc. CVPR, 2023.

Citation

MLA
Koppula, S., et al. “TAPVid-3D: A Benchmark for Tracking Any Point in 3D”. arXiv, 2024, http://arxiv.org/abs/2407.05921v2.
APA
Koppula, S., Rocco, I., Yang, Y., Heyward, J., Carreira, J., Zisserman, A., Brostow, G., & Doersch, C. (2024). TAPVid-3D: A Benchmark for Tracking Any Point in 3D. arXiv. http://arxiv.org/abs/2407.05921v2
Chicago
Koppula, S., I. Rocco, Y. Yang, et al. 2024. “TAPVid-3D: A Benchmark for Tracking Any Point in 3D”. arXiv. http://arxiv.org/abs/2407.05921v2.
Harvard
Koppula, S. et al. (2024) “TAPVid-3D: A Benchmark for Tracking Any Point in 3D”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2407.05921v2.
Vancouver
1. Koppula S, Rocco I, Yang Y, Heyward J, Carreira J, Zisserman A, Brostow G, Doersch C (2024) TAPVid-3D: A Benchmark for Tracking Any Point in 3D. arXiv

BibTeX

@article{koppula2024tapvid,
  title = {TAPVid-3D: A Benchmark for Tracking Any Point in 3D},
  author = {Koppula, Skanda and Rocco, Ignacio and Yang, Yi and Heyward, Joe and Carreira, João and Zisserman, Andrew and Brostow, Gabriel and Doersch, Carl},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2407.05921v2},
  eprint = {2407.05921}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission