Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps
Yue HuShaoheng FangZixing LeiYiqi ZhongSiheng Chen
Proposes Where2comm, a collaborative perception framework that uses spatial confidence maps to transmit only critical perceptual regions, reducing communication bandwidth by over 100,000 times while outperforming state-of-the-art multi-agent 3D detection models.
Autonomous systems operating in complex environments—such as self-driving car fleets, robotic warehouse networks, and drone search-and-rescue teams—often suffer from blind spots, long-range degradation, and severe occlusions. Sharing perceptual data across multiple connected agents significantly improves safety and environmental awareness. However, real-world communication channels have limited, fluctuating bandwidth, creating a critical bottleneck when sharing heavy raw sensory feeds or large feature maps.
The main objective of the article is to develop and evaluate Where2comm, a communication-efficient multi-agent collaborative perception framework. The system is designed to balance perception accuracy and communication bandwidth by dynamically identifying where to communicate, what sparse information to share, and with whom to collaborate across multiple rounds.
To achieve this, the authors designed a framework centered on a spatial confidence generator that produces confidence maps reflecting the perceptual importance of specific areas. Agents transmit compact messages containing only critical, sparse feature areas alongside request maps highlighting regions where they lack information. Feature fusion is managed through a location-specific multi-head attention transformer that incorporates spatial confidence and sensor distance priors. The approach was evaluated on 3D object detection across four datasets: the real-world DAIR-V2X dataset, two vehicle simulation benchmarks (OPV2V and V2X-Sim), and a newly introduced aerial swarm dataset (CoPerception-UAVs) featuring over 131,000 images, covering both camera and LiDAR sensors across vehicles and drones.
The evaluations yielded several significant findings. First, Where2comm established a superior performance-to-bandwidth trade-off across all benchmarks, reducing the required communication volume by factors ranging from 55 to over 100,000 times compared to existing intermediate fusion methods while matching or exceeding their detection accuracy. Second, it improved overall detection performance across benchmarks, raising average precision by 7.7% on real-world vehicle data, 6.62% on drone swarms, and 25.81% on vehicle simulations. Third, multi-round communication consistently boosted accuracy, confirming that targeted requests across rounds yield compounding perceptual benefits. Fourth, the system showed high resilience to realistic localization noise, sustaining performance under spatial position shifts that degraded alternative methods below non-collaborative baselines.
These results demonstrate that multi-agent systems do not require exhaustive data transmission to achieve safe, high-quality collective intelligence. By pruning irrelevant background transmission and focusing purely on perceptually dense regions, organizations can dramatically reduce network operational costs and latency risks. This enables scalable multi-agent deployments even under tight bandwidth constraints and imperfect sensor conditions.
Based on these findings, development teams deploying collaborative autonomous fleets should adopt spatial-confidence-aware message pruning and per-location attention mechanisms. System designers should configure dynamic, multi-round communication protocols rather than single-round broadcasts. To advance deployment readiness, future initiatives should extend this confidence-aware selection strategy to the temporal domain to determine the optimal timing for communication, while validating the framework against transmission latency, asynchronous message arrival, and real-world aerial field tests.
Confidence in these findings is high given the consistent gains across diverse sensors, simulated domains, and real-world traffic data. Nevertheless, leaders should exercise appropriate caution regarding real-world drone deployments, as the aerial evaluations were conducted within high-fidelity simulations that may not fully reflect environmental interference, hardware latency, and communication packet drops encountered in live field operations.
- Paper: Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D, Jonah Philion et al. (2020). It introduces Lift-Splat-Shoot for lifting multi-camera images into unified bird's-eye-view feature maps, which provides the foundational representation used by camera-based collaborative perception frameworks.
- Paper: Center-based 3D Object Detection and Tracking, Tianwei Yin et al. (2020). It establishes point- and center-based 3D object detection representation on bird's-eye-view grids, which is directly relevant to constructing spatial feature maps and confidence scores.
- Paper: PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection, Shaoshuai Shi et al. (2019). It details point-voxel feature integration for 3D point cloud detection, underpinning LiDAR-based multi-agent perception pipelines.
- Paper: Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training, Yujun Lin et al. (2018). It provides foundational principles on extreme sparsity and communication-efficient distributed updates, inspiring sparse multi-agent communication strategies.
- Paper: Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges, Di Feng et al. (2019). It comprehensively reviews multi-modal sensor fusion architectures and datasets for autonomous driving, establishing the core challenges of collaborative and multi-sensor perception.
- Paper: BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation, Zhijian Liu et al. (2022). It introduces a unified multi-task, multi-sensor BEV fusion framework that advances beyond single-modality representations used in early collaborative perception models.
- Paper: BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers, Zhiqi Li et al. (2022). It designs spatiotemporal transformers to learn high-quality bird's-eye-view representations directly from multi-camera sequences, providing an advanced feature encoder for vision-based vehicle perception.
