Argoverse: 3D Tracking and Forecasting With Rich Maps
Ming-Fang ChangJohn LambertPatsorn SangkloyJagjeet SinghSlawomir BakAndrew HartnettDe WangPeter CarrSimon LuceyDeva Ramanan
Introduces a pioneering autonomous driving benchmark that pairs sensor data and stereo imagery with high-definition geometric and semantic maps to advance 3D tracking and motion forecasting.
Autonomous vehicle perception systems require accurate spatial context to safely navigate complex urban environments, yet publicly available benchmarks have historically lacked rich geometric and semantic map data. This limitation has hindered research into how detailed prior map knowledge can assist onboard sensor systems in tasks such as object tracking and path prediction.
The article introduces Argoverse, a large-scale open-source dataset, and evaluates how incorporating high-definition map features directly impacts the accuracy of 3D object tracking and motion forecasting algorithms.
The researchers collected real-world sensor data across Pittsburgh and Miami using autonomous vehicle fleets equipped with long-range LiDAR, 360-degree cameras, and stereo imagery. The resulting benchmark pairs these sensor streams with rich map layers, including 290 linear kilometers of lane centerlines with semantic connectivity, a digital ground height model, and binary driveable area coverage. The dataset features 113 human-annotated vehicle log segments for 3D tracking across 15 object classes and over 324,000 mined five-second scenarios capturing challenging driving maneuvers for trajectory forecasting.
The evaluation established several core findings. First, integrating vector map lane directions substantially improved vehicle orientation estimation during 3D tracking, cutting orientation error nearly in half from roughly 25–28 degrees to 13–15 degrees across near and far ranges. Second, using high-definition maps for ground-point filtering maintained steady 3D bounding box shape accuracy and improved object detection scores compared to traditional planar ground-fitting heuristics, particularly in sloped and uneven terrain. Third, baseline tracking accuracy degraded sharply with distance, dropping from a Multi-Object Tracking Accuracy score of 65.5% within 30 meters down to 34.2% within 100 meters due to sensor point sparsity. Finally, in motion forecasting, leveraging map centerlines as reference priors and using driveable area boundaries to prune invalid paths produced multi-path trajectory forecasts with drivable area compliance rates as high as 94% to 99%.
These findings demonstrate that detailed map priors serve as critical computational aids for perception systems. Incorporating semantic road infrastructure directly addresses false detections and physical violations, enhancing safety and route-planning reliability while reducing onboard real-time processing burdens. The results show that even simple statistical models utilizing map priors can outperform complex, unconstrained machine learning models by restricting predicted behaviors to physically plausible road lanes.
Autonomous driving engineering teams and researchers should adopt multimodal datasets that integrate high-definition vector maps to train perception models. Future technical efforts should prioritize developing advanced machine learning models that jointly leverage camera, LiDAR, and map representations, alongside building automated mapping systems to scale these annotations to new geographic regions.
While the dataset offers substantial scale and diversity, the motion forecasting scenarios rely on automatically mined trajectories from fleet operations, which introduces a degree of label noise compared to manually certified ground truth. Readers should exercise caution when evaluating long-range tracking capabilities, as LiDAR point sparsity beyond 50 meters remains a fundamental hardware constraint.
- Paper: PointRCNN: 3D Object Proposal Generation and Detection From Point Cloud, Shaoshuai Shi et al. (2019). Argoverse's 3D tracking and detection tasks directly rely on foundational 3D proposal generation architectures like PointRCNN that operate on raw LiDAR point clouds.
- Paper: Multi-view 3D Object Detection Network for Autonomous Driving, Xiaozhi Chen et al. (2017). Argoverse evaluates multi-sensor 3D perception baselines that build on multi-view fusion architectures combining LiDAR point clouds and camera imagery pioneered by MV3D.
- Paper: MOT16: A Benchmark for Multi-Object Tracking, Anton Milan et al. (2016). Argoverse adapts standardized multi-object tracking metrics and benchmark evaluation formulations established by benchmarks like MOT16 to the 3D domain.
- Paper: The Cityscapes Dataset for Semantic Urban Scene Understanding, Marius Cordts et al. (2016). Cityscapes established the paradigm of large-scale, multi-city urban datasets for autonomous vehicle scene understanding that Argoverse expands with LiDAR and HD maps.
- Paper: Object scene flow for autonomous vehicles, Moritz Menze et al. (2015). Menze and Geiger established foundational principles and dynamic scene benchmarks for multi-object 3D motion estimation from stereo imagery in autonomous driving.
- Paper: Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D, Jonah Philion et al. (2020). Lift, Splat, Shoot builds upon modern multi-camera autonomous driving benchmark suites like Argoverse to learn unified bird's-eye-view representations for downstream planning and forecasting.
- Paper: BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers, Zhiqi Li et al. (2022). BEVFormer advances the multi-camera 3D detection and tracking problems framed by Argoverse by incorporating spatiotemporal transformers in bird's-eye view.
- Paper: BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation, Zhijian Liu et al. (2022). BEVFusion extends multi-sensor 3D perception on urban AV datasets by developing an optimized shared bird's-eye-view representation for LiDAR and camera fusion.
- Paper: Center-based 3D Object Detection and Tracking, Tianwei Yin et al. (2020). CenterPoint introduces point-centered 3D detection and velocity-based tracking methods tailored to large-scale urban driving benchmarks.
- Paper: BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning, Fisher Yu et al. (2020). BDD100K expands the scale and multi-task scope of driving benchmarks beyond the tracking and trajectory forecasting tasks featured in Argoverse.
- Paper: Adaptive Graph Convolutional Recurrent Network for Traffic Forecasting, Lei Bai et al. (2020). Adaptive Graph Convolutional Recurrent Network builds on the spatio-temporal road graph forecasting challenges introduced in autonomous driving datasets.
