Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps

Yue HuShaoheng FangZixing LeiYiqi ZhongSiheng Chen

article2022NeurIPS412 citationsSpotlight

Proposes Where2comm, a collaborative perception framework that uses spatial confidence maps to transmit only critical perceptual regions, reducing communication bandwidth by over 100,000 times while outperforming state-of-the-art multi-agent 3D detection models.

Listen

Autonomous systems operating in complex environments—such as self-driving car fleets, robotic warehouse networks, and drone search-and-rescue teams—often suffer from blind spots, long-range degradation, and severe occlusions. Sharing perceptual data across multiple connected agents significantly improves safety and environmental awareness. However, real-world communication channels have limited, fluctuating bandwidth, creating a critical bottleneck when sharing heavy raw sensory feeds or large feature maps.

The main objective of the article is to develop and evaluate Where2comm, a communication-efficient multi-agent collaborative perception framework. The system is designed to balance perception accuracy and communication bandwidth by dynamically identifying where to communicate, what sparse information to share, and with whom to collaborate across multiple rounds.

To achieve this, the authors designed a framework centered on a spatial confidence generator that produces confidence maps reflecting the perceptual importance of specific areas. Agents transmit compact messages containing only critical, sparse feature areas alongside request maps highlighting regions where they lack information. Feature fusion is managed through a location-specific multi-head attention transformer that incorporates spatial confidence and sensor distance priors. The approach was evaluated on 3D object detection across four datasets: the real-world DAIR-V2X dataset, two vehicle simulation benchmarks (OPV2V and V2X-Sim), and a newly introduced aerial swarm dataset (CoPerception-UAVs) featuring over 131,000 images, covering both camera and LiDAR sensors across vehicles and drones.

The evaluations yielded several significant findings. First, Where2comm established a superior performance-to-bandwidth trade-off across all benchmarks, reducing the required communication volume by factors ranging from 55 to over 100,000 times compared to existing intermediate fusion methods while matching or exceeding their detection accuracy. Second, it improved overall detection performance across benchmarks, raising average precision by 7.7% on real-world vehicle data, 6.62% on drone swarms, and 25.81% on vehicle simulations. Third, multi-round communication consistently boosted accuracy, confirming that targeted requests across rounds yield compounding perceptual benefits. Fourth, the system showed high resilience to realistic localization noise, sustaining performance under spatial position shifts that degraded alternative methods below non-collaborative baselines.

These results demonstrate that multi-agent systems do not require exhaustive data transmission to achieve safe, high-quality collective intelligence. By pruning irrelevant background transmission and focusing purely on perceptually dense regions, organizations can dramatically reduce network operational costs and latency risks. This enables scalable multi-agent deployments even under tight bandwidth constraints and imperfect sensor conditions.

Based on these findings, development teams deploying collaborative autonomous fleets should adopt spatial-confidence-aware message pruning and per-location attention mechanisms. System designers should configure dynamic, multi-round communication protocols rather than single-round broadcasts. To advance deployment readiness, future initiatives should extend this confidence-aware selection strategy to the temporal domain to determine the optimal timing for communication, while validating the framework against transmission latency, asynchronous message arrival, and real-world aerial field tests.

Confidence in these findings is high given the consistent gains across diverse sensors, simulated domains, and real-world traffic data. Nevertheless, leaders should exercise appropriate caution regarding real-world drone deployments, as the aerial evaluations were conducted within high-fidelity simulations that may not fully reflect environmental interference, hardware latency, and communication packet drops encountered in live field operations.

Cover for Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps

Abstract

Multi-agent collaborative perception could significantly upgrade the perception performance by enabling agents to share complementary information with each other through communication. It inevitably results in a fundamental trade-off between perception performance and communication bandwidth. To tackle this bottleneck issue, we propose a spatial confidence map, which reflects the spatial heterogeneity of perceptual information. It empowers agents to only share spatially sparse, yet perceptually critical information, contributing to where to communicate. Based on this novel spatial confidence map, we propose Where2comm, a communication-efficient collaborative perception framework. Where2comm has two distinct advantages: i) it considers pragmatic compression and uses less communication to achieve higher perception performance by focusing on perceptually critical areas; and ii) it can handle varying communication bandwidth by dynamically adjusting spatial areas involved in communication. To evaluate Where2comm, we consider 3D object detection in both real-world and simulation scenarios with two modalities (camera/LiDAR) and two agent types (cars/drones) on four datasets: OPV2V, V2X-Sim, DAIR-V2X, and our original CoPerception-UAVs. Where2comm consistently outperforms previous methods; for example, it achieves more than 100,000×100,000 \times lower communication volume and still outperforms DiscoNet and V2X-ViT on OPV2V. Our code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Problem Formulation
  • 4 Where2comm: Spatial Confidence-Aware Collaborative Perception System
  • 4.1 Observation encoder
  • 4.2 Spatial confidence generator
  • 4.3 Spatial confidence-aware communication
  • 4.4 Spatial confidence-aware message fusion
  • 4.5 Detection decoder
  • 4.6 Training details and loss functions
  • 5 Experimental Results
  • 5.1 Datasets and experimental settings
  • 5.2 Quantitative evaluation
  • 5.3 Qualitative evaluation
  • 5.4 Ablation studies
  • 6 Conclusion and limitation
  • References
  • 7 Appendix
  • 7.1 Highlights of our contribution
  • 7.2 Detailed information about the system pipeline
  • 7.3 Detailed information about the optimization problem of collaborative perception
  • 7.4 Detailed information about the module design
  • 7.5 Detailed information about experimental settings
  • 7.6 Benchmarks
  • 7.7 Visualization
  • 7.8 Ablation on bandwidth allocation
  • 7.9 Discussion on the realistic limitations
  • 7.10 CoPerception-UAVs dataset details

Knowls

  1. Knowl 1 — Multi-Round Spatial Confidence-Aware Collaborative Perception Algorithm

    algorithm

    Where2comm executes multi-agent collaborative perception across KK communication rounds by iteratively generating spatial confidence maps, exchanging spatially sparse features guided by confidence and request priors, and fusing received features using confidence-aware attention.

    Input: Number of agents NN, total communication rounds KK, sensory observations {Xi}i=1N\{X_i\}_{i=1}^N
    Output: Multi-agent 3D bounding box detection predictions {O^i(K)}i=1N\{\hat{O}_i^{(K)}\}_{i=1}^N
    for i=1i = 1 to NN do
        Fi(0)←Φenc(Xi)∈RH×W×DF_i^{(0)} \leftarrow \Phi_{\text{enc}}(X_i) \in \mathbb{R}^{H \times W \times D}
    end for
    for k=0k = 0 to K−1K - 1 do
        for i=1i = 1 to NN do
            Ci(k)←Φgenerator(Fi(k))∈[0,1]H×WC_i^{(k)} \leftarrow \Phi_{\text{generator}}(F_i^{(k)}) \in [0, 1]^{H \times W}
            Ri(k)←1−Ci(k)∈[0,1]H×WR_i^{(k)} \leftarrow 1 - C_i^{(k)} \in [0, 1]^{H \times W}
            for j=1j = 1 to NN do
                if k=0k = 0 then
                    Mi→j(0)←Φselect(Ci(0))∈{0,1}H×WM_{i \to j}^{(0)} \leftarrow \Phi_{\text{select}}(C_i^{(0)}) \in \{0, 1\}^{H \times W}
                    Ai,j(0)←1A_{i, j}^{(0)} \leftarrow 1
                else
                    Mi→j(k)←Φselect(Ci(k)⊙Rj(k−1))∈{0,1}H×WM_{i \to j}^{(k)} \leftarrow \Phi_{\text{select}}(C_i^{(k)} \odot R_j^{(k-1)}) \in \{0, 1\}^{H \times W}
                    Ai,j(k)←max⁡h,w(Mi→j(k))h,w∈{0,1}A_{i, j}^{(k)} \leftarrow \max_{h, w} (M_{i \to j}^{(k)})_{h, w} \in \{0, 1\}
                end if
                Zi→j(k)←Mi→j(k)⊙Fi(k)∈RH×W×DZ_{i \to j}^{(k)} \leftarrow M_{i \to j}^{(k)} \odot F_i^{(k)} \in \mathbb{R}^{H \times W \times D}
            end for
        end for
        for i=1i = 1 to NN do
            for j=1j = 1 to NN with j≠ij \neq i and Ai,j(k)=1A_{i, j}^{(k)} = 1 do
                Transmit message Pi→j(k)=(Zi→j(k),Ri(k))P_{i \to j}^{(k)} = (Z_{i \to j}^{(k)}, R_i^{(k)}) to agent jj
            end for
            Receive incoming messages {Pj→i(k)=(Zj→i(k),Rj(k)):j≠i,Aj,i(k)=1}\{P_{j \to i}^{(k)} = (Z_{j \to i}^{(k)}, R_j^{(k)}) : j \neq i, A_{j, i}^{(k)} = 1\}
            Fi(k+1)←ffuse(Fi(k),{Pj→i(k)})∈RH×W×DF_i^{(k+1)} \leftarrow f_{\text{fuse}}\left(F_i^{(k)}, \{P_{j \to i}^{(k)}\}\right) \in \mathbb{R}^{H \times W \times D}
        end for
    end for
    for i=1i = 1 to NN do
        O^i(K)←Φdec(Fi(K))\hat{O}_i^{(K)} \leftarrow \Phi_{\text{dec}}(F_i^{(K)})
    end for
    return {O^i(K)}i=1N\{\hat{O}_i^{(K)}\}_{i=1}^N

    In the algorithm, Φenc\Phi_{\text{enc}} extracts Bird's Eye View (BEV) features of dimensions H×W×DH \times W \times D from camera or LiDAR data. The spatial confidence generator Φgenerator\Phi_{\text{generator}} outputs confidence map Ci(k)C_i^{(k)}, which yields request map Ri(k)R_i^{(k)}. Selection function Φselect\Phi_{\text{select}} selects top spatial locations meeting bandwidth limits, creating selection mask Mi→j(k)M_{i \to j}^{(k)} and sparsified feature map Zi→j(k)Z_{i \to j}^{(k)}. Adjacency matrix A(k)A^{(k)} defines the communication graph. Message fusion ffusef_{\text{fuse}} fuses sparse incoming features, and Φdec\Phi_{\text{dec}} outputs bounding box predictions O^i(K)∈RH×W×7\hat{O}_i^{(K)} \in \mathbb{R}^{H \times W \times 7} containing position (x,y)(x, y), size (h,w)(h, w), heading (cos⁡α,sin⁡α)(\cos \alpha, \sin \alpha), and classification confidence cc.

  2. Knowl 2 — Spatial Confidence Map and Request Map Generation

    model/method

    In Where2comm, perceptual information exhibits spatial heterogeneity across a shared Bird's Eye View (BEV) coordinate grid of dimensions H×WH \times W. For agent ii at communication round kk, with intermediate feature representation Fi(k)∈RH×W×DF_i^{(k)} \in \mathbb{R}^{H \times W \times D}, a spatial confidence map Ci(k)C_i^{(k)} is produced via:

    Ci(k)=Φgenerator(Fi(k))∈[0,1]H×WC_i^{(k)} = \Phi_{\text{generator}}\left(F_i^{(k)}\right) \in [0, 1]^{H \times W}

    where Φgenerator(⋅)\Phi_{\text{generator}}(\cdot) shares the architecture and weights of the classification head in the object detection decoder Φdec(⋅)\Phi_{\text{dec}}(\cdot) to ensure parameter efficiency. Each value in Ci(k)C_i^{(k)} quantifies the confidence that a foreground object exists at that spatial grid cell.

    The corresponding request map Ri(k)R_i^{(k)} represents areas of perceptual uncertainty or missing observation (such as occluded or out-of-range regions) and is defined as the complement of the spatial confidence map:

    Ri(k)=1−Ci(k)∈[0,1]H×WR_i^{(k)} = 1 - C_i^{(k)} \in [0, 1]^{H \times W}

    High values in Ri(k)R_i^{(k)} explicitly indicate spatial regions where agent ii requires complementary perceptual information from neighboring agents in subsequent rounds.

  3. Knowl 3 — Spatial Confidence-Aware Message Packing and Communication Graph Construction

    model/method

    Where2comm constructs communication messages and peer connectivity dynamically based on spatial confidence and request maps to achieve bandwidth-efficient collaboration.

    For a message sent from agent ii to agent jj at communication round kk, a binary spatial selection matrix Mi→j(k)∈{0,1}H×WM_{i \to j}^{(k)} \in \{0, 1\}^{H \times W} is computed via:

    Mi→j(k)={Φselect(Ci(k)),k=0Φselect(Ci(k)⊙Rj(k−1)),k>0M_{i \to j}^{(k)} = \begin{cases} \Phi_{\text{select}}\left(C_i^{(k)}\right), & k = 0 \\ \Phi_{\text{select}}\left(C_i^{(k)} \odot R_j^{(k-1)}\right), & k > 0 \end{cases}

    where ⊙\odot denotes the element-wise Hadamard product, Ci(k)C_i^{(k)} is agent ii's confidence map, Rj(k−1)R_j^{(k-1)} is the request map received from agent jj in the previous round, and Φselect(⋅)\Phi_{\text{select}}(\cdot) sets the top spatial locations ranking highest to 11 up to the communication budget constraint (optionally filtered with a 2D Gaussian kernel to incorporate neighborhood context and suppress isolated noise).

    The transmitted feature map is sparsified as Zi→j(k)=Mi→j(k)⊙Fi(k)∈RH×W×DZ_{i \to j}^{(k)} = M_{i \to j}^{(k)} \odot F_i^{(k)} \in \mathbb{R}^{H \times W \times D}. The transmitted message is Pi→j(k)=(Ri(k),Zi→j(k))P_{i \to j}^{(k)} = (R_i^{(k)}, Z_{i \to j}^{(k)}), transmitting only non-zero feature vectors and their corresponding spatial indices.

    The communication graph adjacency matrix A(k)∈{0,1}N×NA^{(k)} \in \{0, 1\}^{N \times N} is constructed per-location via:

    Ai,j(k)={1,k=0max⁡h∈{0,…,H−1},w∈{0,…,W−1}(Mi→j(k))h,w,k>0A_{i, j}^{(k)} = \begin{cases} 1, & k = 0 \\ \max_{h \in \{0, \dots, H-1\}, w \in \{0, \dots, W-1\}} \left(M_{i \to j}^{(k)}\right)_{h, w}, & k > 0 \end{cases}

    When Mi→j(k)M_{i \to j}^{(k)} contains all zeros, communication between agent ii and agent jj is completely pruned. The communication volume for Zi→j(k)Z_{i \to j}^{(k)} in base-2 logarithmic bytes is computed as:

    Communication Volume=log⁡2(∣Mi→j(k)∣0×D×328)\text{Communication Volume} = \log_2\left(|M_{i \to j}^{(k)}|_0 \times D \times \frac{32}{8}\right)

    where ∣⋅∣0|\cdot|_0 counts the number of selected non-zero spatial cells, DD is the channel dimension, 32 represents bits per float32 value, and division by 8 yields bytes.

  4. Knowl 4 — Spatial Confidence-Aware Per-Location Message Fusion

    model/method

    Where2comm fuses multi-agent intermediate representations using a transformer architecture operating per BEV grid location, incorporating sender spatial confidence maps and physical sensing distances as explicit attention priors.

    Upon receiving message Pj→i(k)=(Zj→i(k),Rj(k))P_{j \to i}^{(k)} = (Z_{j \to i}^{(k)}, R_j^{(k)}), agent ii recovers the sender confidence map Cj(k)=1−Rj(k)C_j^{(k)} = 1 - R_j^{(k)}. Ego features are denoted Zi→i(k)=Fi(k)Z_{i \to i}^{(k)} = F_i^{(k)} with confidence Ci(k)C_i^{(k)}. The per-location cross-agent attention weight matrix Wj→i(k)∈RH×WW_{j \to i}^{(k)} \in \mathbb{R}^{H \times W} is calculated as:

    Wj→i(k)=MHAW(Fi(k),Zj→i(k),Zj→i(k))⊙Cj(k)W_{j \to i}^{(k)} = \text{MHAW}\left(F_i^{(k)}, Z_{j \to i}^{(k)}, Z_{j \to i}^{(k)}\right) \odot C_j^{(k)}

    where MHAW(⋅)\text{MHAW}(\cdot) denotes scaled dot-product multi-head attention applied independently at each spatial cell (h,w)(h, w). The updated feature representation for agent ii is:

    Fi(k+1)=FFN(∑j∈Ni∪{i}Wj→i(k)⊙Zj→i(k))∈RH×W×DF_i^{(k+1)} = \text{FFN}\left(\sum_{j \in \mathcal{N}_i \cup \{i\}} W_{j \to i}^{(k)} \odot Z_{j \to i}^{(k)}\right) \in \mathbb{R}^{H \times W \times D}

    where FFN(⋅)\text{FFN}(\cdot) is a feed-forward network and Ni={j∣Aj,i(k)=1,j≠i}\mathcal{N}_i = \{j \mid A_{j, i}^{(k)} = 1, j \neq i\} is the set of transmitting neighbors.

    Before transformer fusion, sensor positional encoding (SPE) is added to BEV feature tokens based on Euclidean distance disdis in 3D physical space between the sensor origin and the grid cell (h,w)(h, w):

    SPE(dis,2p)=sin⁡(dis100002p/D),SPE(dis,2p+1)=cos⁡(dis100002p/D)\text{SPE}(dis, 2p) = \sin\left(\frac{dis}{10000^{2p/D}}\right), \quad \text{SPE}(dis, 2p+1) = \cos\left(\frac{dis}{10000^{2p/D}}\right)

    where p∈{0,…,D/2−1}p \in \{0, \dots, D/2 - 1\} indexes the channel dimension and DD is the total feature channel dimension.

  5. Knowl 5 — Optimization Decomposition for Bandwidth-Constrained Collaborative Perception

    model/method

    Collaborative perception is formulated as a constrained optimization problem maximizing collective perception performance g(⋅)g(\cdot) over NN agents under a total communication budget B1B_1:

    max⁡θ,P∑i=1Ng(Φθ(Xi,{Pi→j}j=1N),Yi)s.t.∑i=1N∑j=1,j≠iN∣Pi→j∣≤B1\max_{\theta, P} \sum_{i=1}^N g\left(\Phi_\theta\left(X_i, \{P_{i \to j}\}_{j=1}^N\right), Y_i\right) \quad \text{s.t.} \quad \sum_{i=1}^N \sum_{j=1, j \neq i}^N |P_{i \to j}| \le B_1

    where XiX_i is the observation, YiY_i is the ground-truth annotation of agent ii, Φθ\Phi_\theta is the perception network parameterized by θ\theta, and Pi→j=Mi→j⊙FiP_{i \to j} = M_{i \to j} \odot F_i with binary selection matrix Mi→j∈{0,1}H×WM_{i \to j} \in \{0, 1\}^{H \times W}. Direct end-to-end optimization of this objective is intractable due to the non-differentiable binary constraint ∑i,j∣Mi→j∣≤b1=B1/D\sum_{i, j} |M_{i \to j}| \le b_1 = B_1 / D.

    Where2comm decomposes the problem into two decoupled sub-optimization stages:

    1. Feasible selection matrix generation via a proxy constrained objective: max⁡M∑i=1N∑j=1,j≠iNMi→j⊙Cis.t.∑i=1N∑j=1,j≠iN∣Mi→j∣0≤b1,Mi→j∈{0,1}H×W\max_{M} \sum_{i=1}^N \sum_{j=1, j \neq i}^N M_{i \to j} \odot C_i \quad \text{s.t.} \quad \sum_{i=1}^N \sum_{j=1, j \neq i}^N |M_{i \to j}|_0 \le b_1, \quad M_{i \to j} \in \{0, 1\}^{H \times W} where CiC_i is the spatial confidence map. This sub-problem has an exact analytical solution obtained by selecting the top-b1b_1 ranked spatial locations in CiC_i.

    2. Network parameter optimization: max⁡θ∑i=1Ng(Φθ(Xi,{Mi→j⊙Fi}j=1N),Yi)\max_\theta \sum_{i=1}^N g\left(\Phi_\theta\left(X_i, \{M_{i \to j} \odot F_i\}_{j=1}^N\right), Y_i\right) With Mi→jM_{i \to j} determined analytically, θ\theta is optimized unconstrained using standard supervised backpropagation on detection loss L=∑k=0K∑i=1NLdet(O^i(k),Oi)L = \sum_{k=0}^K \sum_{i=1}^N L_{\text{det}}(\hat{O}_i^{(k)}, O_i).

  6. Knowl 6 — CoPerception-UAVs Dataset for Multi-UAV Collaborative Perception

    experimental setup

    CoPerception-UAVs is a collaborative perception dataset simulated using CARLA and AirSim to benchmark multi-agent perception in unmanned aerial vehicle (UAV) swarms.

    Dataset specifications:

    • Total images and 3D boxes: 131.9k synchronous images and 1.94M 3D bounding boxes (with 3.6M 2D bounding boxes and pixel-wise/BEV semantic masks).
    • Data split: 91,175 training images (1,316,536 3D boxes), 19,500 validation images (303,888 3D boxes), and 20,250 testing images (319,576 3D boxes), structured according to nuScenes-devkit format.
    • Swarm composition: 5 coordinated UAVs flying across 3 altitudes over three CARLA urban maps (town4, town5, town6) with an operational perception range of 200 m×350 m200\,\text{m} \times 350\,\text{m}.
    • Sensor configuration per UAV: 5 RGB cameras and 5 semantic cameras (1 downward BEV camera, 4 directional cameras at 90∘90^\circ FoV, −45∘-45^\circ pitch, resolution 800×450800 \times 450), covering approximately 200 m×200 m200\,\text{m} \times 200\,\text{m} at 40 m40\,\text{m} altitude.
    • Swarm flight formations:
      1. Discipline mode (123.8k images): UAVs maintain a fixed, structured array formation (modeling area exploration / search-and-rescue).
      2. Dynamic mode (8.1k images): Each UAV navigates independently (modeling urban surveillance / patrol).
  7. Knowl 7 — Benchmark Comparison of Collaborative 3D Object Detection Performance

    data/table

    Where2comm was evaluated against collaborative perception baselines across four datasets: CoPerception-UAVs (camera-only aerial detection), OPV2V (camera-only autonomous driving detection), V2X-Sim 1.0 / 2.0 (LiDAR-based vehicle-to-everything detection), and DAIR-V2X (real-world vehicle-infrastructure LiDAR detection). Detection performance is measured by Average Precision at IoU thresholds of 0.50 and 0.70 ([email protected] / [email protected]), and communication volume is measured in log⁡2(bytes)\log_2(\text{bytes}) per message.

    Dataset CoPerception-UAVs OPV2V V2X-Sim1.0 DAIR-V2X
    Method Comm [email protected]/0.70 Comm [email protected]/0.70 [email protected] Comm [email protected]/0.70
    No Collaboration 0.00 57.67 / 29.52 0.00 22.65 / 9.09 45.80 0.00 50.03 / 43.57
    Late Fusion 15.77 53.12 / 37.88 11.87 8.24 / 3.84 46.70 11.45 53.12 / 37.88
    When2com 28.37 61.63 / 33.55 22.28 19.69 / 8.29 46.70 22.62 51.12 / 36.17
    V2VNet 29.95 59.82 / 33.14 23.87 37.47 / 14.67 55.30 24.21 56.01 / 42.25
    V2X-ViT 28.37 59.12 / 41.57 22.28 39.82 / 16.43 57.30 22.62 54.26 / 43.35
    DiscoNet 28.37 59.74 / 29.71 22.28 36.00 / 12.50 58.00 22.62 54.29 / 44.88
    Where2comm (low) 11.76 60.19 / 34.94 5.67 40.11 / 15.36 47.60 11.40 50.98 / 39.11
    Where2comm (mid) 17.96 63.04 / 36.10 17.04 44.07 / 17.15 51.80 17.53 55.84 / 42.44
    Where2comm (high) 28.48 65.71 / 39.38 22.71 47.14 / 19.07 59.10 22.62 63.71 / 48.93

    On V2X-Sim 2.0 (LiDAR), Where2comm achieves 75.72/65.13 at Comm=13.84 and 83.77/74.09 at Comm=25.93, outperforming DiscoNet (69.73/55.12 at Comm=26.04), V2X-ViT (78.73/63.17 at Comm=26.04), and V2VNet (80.80/71.22 at Comm=27.62).

    Across all benchmarks, Where2comm achieves higher detection performance with substantially lower communication bandwidth (e.g., matching or exceeding prior state-of-the-art methods with over 100,000×100{,}000\times less communication volume on OPV2V, 5,128×5{,}128\times less on CoPerception-UAVs, 105×105\times less on DAIR-V2X, and 55×55\times less on V2X-Sim).

  8. Knowl 8 — Multi-Round Communication Dynamics and Bandwidth Allocation Ratios

    empirical result

    Evaluating Where2comm across varying communication rounds (K∈{1,2,3}K \in \{1, 2, 3\}) and bandwidth allocation splits reveals two key properties:

    1. Multi-round scalability: Increasing communication rounds from 1 to 3 steadily shifts the performance-bandwidth trade-off curve upward across CoPerception-UAVs, OPV2V, and V2X-Sim datasets, providing consistently higher [email protected] at equivalent total bandwidth budgets.

    2. Targeted bandwidth allocation: In multi-round communication, allocating a smaller bandwidth proportion to the initial broadcast round (round 0) and the majority of bandwidth to subsequent rounds yields superior performance compared to allocating all bandwidth upfront (ratio 1.0−01.0 - 0). For a 2-round setting, an allocation ratio of 0.2−0.80.2 - 0.8 (20% broadcast, 80% request-response) outperforms 0.5−0.50.5 - 0.5 and 0.8−0.20.8 - 0.2. For a 3-round setting, an allocation ratio of 0.2−0.6−0.20.2 - 0.6 - 0.2 achieves the best trade-off. This occurs because rounds k>0k > 0 use request maps (Rj(k−1)R_j^{(k-1)}) from collaborators to transmit only the spatial features requested by peers.

  9. Knowl 9 — Ablation of Message Fusion Components and Selection Filtering

    data/table

    Ablation experiments evaluate the contribution of Multi-Head Attention (MHA), Sensor Positional Encoding (SPE), and Spatial Confidence Maps (SCM) in the message fusion module across OPV2V, CoPerception-UAVs, and V2X-Sim datasets.

    MHA SPE SCM OPV2V ([email protected]/0.70) CoPerception-UAVs ([email protected]/0.70) V2X-Sim ([email protected]/0.70)
    – – – 34.96 / 13.92 63.48 / 44.23 51.2 / 45.7
    ✓ – – 38.75 / 13.28 63.99 / 44.46 57.3 / 50.8
    ✓ ✓ – 39.82 / 16.43 64.34 / 46.86 59.1 / 52.0
    ✓ ✓ ✓ 47.30 / 19.30 64.83 / 47.62 59.1 / 52.2
    • Multi-head attention (MHA) improves [email protected] over vanilla fusion by 3.79% on OPV2V and 6.1% on V2X-Sim by modeling fine-grained per-cell interactions across multiple attention heads.
    • Sensor Positional Encoding (SPE) and Spatial Confidence Maps (SCM) act as complementary priors, with the full configuration improving [email protected]/0.70 on OPV2V from 34.96/13.92 to 47.30/19.30 (+12.34% / +5.38%).
    • Adding a 2D Gaussian filter during binary selection matrix computation (Mi→j(k)M_{i \to j}^{(k)}) consistently improves detection AP across communication bandwidth levels by smoothing isolated noise and preserving local spatial context.
  10. Knowl 10 — Robustness of Collaborative Perception to Sensor Localization Noise

    empirical result

    The robustness of Where2comm against coordinate localization error was evaluated by adding zero-mean Gaussian noise with standard deviation ranging from 0 m0\,\text{m} to 0.6 m0.6\,\text{m} to agent pose estimates on CoPerception-UAVs, OPV2V, and V2X-Sim.

    Key findings:

    • Where2comm consistently maintains higher 3D object detection [email protected] than baseline methods (When2com, V2VNet, DiscoNet) across all noise levels.
    • On CoPerception-UAVs, baseline collaboration degrades below the single-agent "No Collaboration" level (57.67% [email protected]) when the noise standard deviation exceeds 0.4 m0.4\,\text{m} for V2VNet and 0.5 m0.5\,\text{m} for DiscoNet. In contrast, Where2comm remains superior to No Collaboration across the entire tested range up to 0.6 m0.6\,\text{m}.
    • This robustness stems from the per-location transformer fusion mechanism selectively attending to reliable feature locations and the spatial confidence map suppressing misaligned, low-confidence background regions.
  11. Knowl 11 — Limitations in Temporal Dynamics and Latency Modeling

    limitation

    Where2comm focuses exclusively on optimizing spatial feature selection (determining where to communicate within synchronous Bird's Eye View representations). It does not optimize temporal communication scheduling (determining when across continuous time frames agents should transmit data). Additionally, the framework lacks explicit motion forecasting or delay-compensation modules to handle asynchronous transmission latencies and dropped packets in real-world wireless channels.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Tsun-Hsuan Wang, Sivabalan Manivasagam, Ming Liang, Bin Yang, Wenyuan Zeng, and Raquel Urtasun. V2vnet: Vehicle-to-vehicle communication for joint perception and prediction. In European Conference on Computer Vision, pages 605–621. Springer, 2020.
  2. 2.Yiming Li, Shunli Ren, Pengxiang Wu, Siheng Chen, Chen Feng, and Wenjun Zhang. Learning distilled collaboration graph for multi-agent perception. Advances in Neural Information Processing Systems, 34, 2021.
  3. 3.Siheng Chen, Baoan Liu, Chen Feng, Carlos Vallespi-Gonzalez, and Carl K. Wellington. 3d point cloud processing and learning for autonomous driving: Impacting map creation, localization, and perception. IEEE Signal Processing Magazine, 38:68–86, 2021.
  4. 4.Zhi Li, Ali Vatankhah Barenji, Jiazhi Jiang, Ray Y Zhong, and Gangyan Xu. A mechanism for scheduling multi robot intelligent warehouse system face with dynamic demand. Journal of Intelligent Manufacturing, 31(2):469–480, 2020.
  5. 5.Michela Zaccaria, Mikhail Giorgini, Riccardo Monica, and Jacopo Aleotti. Multi-robot multiple camera people detection and tracking in automated warehouses. In 2021 IEEE 19th International Conference on Industrial Informatics (INDIN), pages 1–6. IEEE, 2021.
  6. 6.Jürgen Scherer, Saeed Yahyanejad, Samira Hayat, Evsen Yanmaz, Torsten Andre, Asif Khan, Vladimir Vukadinovic, Christian Bettstetter, Hermann Hellwagner, and Bernhard Rinner. An autonomous multi-uav system for search and rescue. In Proceedings of the First Workshop on Micro Aerial Vehicle Networks, Systems, and Applications for Civilian Use, pages 33–38, 2015.
  7. 7.Ebtehal Turki Alotaibi, Shahad Saleh Alqefari, and Anis Koubaa. Lsar: Multi-uav collaboration for search and rescue missions. IEEE Access, 7:55817–55832, 2019.
  8. 8.Yue Hu, Shaoheng Fang, Weidi Xie, and Siheng Chen. Aerial monocular 3d object detection. arXiv preprint arXiv:2208.03974, 2022.
  9. 9.Yiming Li, Ziyan An, Zixun Wang, Yiqi Zhong, Siheng Chen, and Chen Feng. V2X-Sim: A virtual collaborative perception dataset for autonomous driving. IEEE Robotics and Automation Letters, 7, 2022.
  10. 10.Runsheng Xu, Hao Xiang, Xin Xia, Xu Han, Jinlong Liu, and Jiaqi Ma. OPV2V: An open benchmark dataset and fusion pipeline for perception with vehicle-to-vehicle communication. ICRA, 2022.
  11. 11.Haibao Yu, Yizhen Luo, Mao Shu, Yiyi Huo, Zebang Yang, Yifeng Shi, Zhenglong Guo, Hanyu Li, Xing Hu, Jirui Yuan, et al. DAIR-V2X: A large-scale dataset for vehicle-infrastructure cooperative 3d object detection. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition (CVPR), 2022.
  12. 12.Yen-Cheng Liu, Junjiao Tian, Nathaniel Glaser, and Zsolt Kira. When2com: Multi-agent perception via communication graph grouping. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pages 4106–4115, 2020.
  13. 13.Yen-Cheng Liu, Junjiao Tian, Chih-Yao Ma, Nathan Glaser, Chia-Wen Kuo, and Zsolt Kira. Who2com: Collaborative perception via learnable handshake communication. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 6876–6883. IEEE, 2020.
  14. 14.Yang Zhou, Jiuhong Xiao, Yue Zhou, and Giuseppe Loianno. Multi-robot collaborative perception with graph neural networks. IEEE Robotics and Automation Letters, 2022.
  15. 15.Eduardo Arnold, Mehrdad Dianati, and Robert de Temple. Cooperative perception for 3d object detection in driving scenarios using infrastructure sensors. IEEE Transactions on Intelligent Transportation Systems, 23:1852–1864, 2022.
  16. 16.Zixing Lei, Shunli Ren, Yue Hu, Wenjun Zhang, and Siheng Chen. Latency-aware collaborative perception. ECCV, 2022.
  17. 17.Yiming Li, Juexiao Zhang, Dekun Ma, Yue Wang, and Chen Feng. Multi-robot scene completion: Towards task-agnostic collaborative perception. In Conference on Robot Learning (CoRL). PMLR, 2022.
  18. 18.Runsheng Xu, Zhengzhong Tu, Hao Xiang, Wei Shao, Bolei Zhou, and Jiaqi Ma. CoBEVT: Cooperative bird’s eye view semantic segmentation with sparse transformers. CoRL, 2022.
  19. 19.Sanbao Su, Yiming Li, Sihong He, Songyang Han, Chen Feng, Caiwen Ding, and Fei Miao. Uncertainty quantification of collaborative detection for self-driving, 2022.
  20. 20.Amanpreet Singh, Tushar Jain, and Sainbayar Sukhbaatar. Learning when to communicate at scale in multiagent cooperative and competitive tasks. ICLR, 2019.
  21. 21.Ming Tan. Multi-agent reinforcement learning: Independent vs. cooperative agents. In Proceedings of the tenth international conference on machine learning, pages 330–337, 1993.
  22. 22.Faisal Qureshi and Demetri Terzopoulos. Smart camera networks in virtual reality. Proceedings of the IEEE, 96(10):1640–1656, 2008.
  23. 23.Yiming Li, Bir Bhanu, and Wei Lin. Auction protocol for camera active control. In 2010 IEEE International Conference on Image Processing, pages 4325–4328. IEEE, 2010.
  24. 24.Sainbayar Sukhbaatar, Rob Fergus, et al. Learning multiagent communication with backpropagation. Advances in neural information processing systems, 29, 2016.
  25. 25.Yedid Hoshen. Vain: Attentional multi-agent predictive modeling. Advances in Neural Information Processing Systems, 30, 2017.
  26. 26.Runsheng Xu, Hao Xiang, Zhengzhong Tu, Xin Xia, Ming-Hsuan Yang, and Jiaqi Ma. V2X-ViT: Vehicle-to-everything cooperative perception with vision transformer. ECCV, 2022.
  27. 27.Y Yuan and M Sester. Comap: A synthetic dataset for collective multi-agent perception of autonomous driving. The International Archives of Photogrammetry, Remote Sensing and Spatial Information Sciences, 43:255–263, 2021.
  28. 28.Xingyi Zhou, Dequan Wang, and Philipp Krähenbühl. Objects as points. In arXiv preprint arXiv:1904.07850, 2019.
  29. 29.Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning, pages 41–48, 2009.
  30. 30.Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. Carla: An open urban driving simulator. In Conference on robot learning, pages 1–16. PMLR, 2017.
  31. 31.Cody Reading, Ali Harakeh, Julia Chae, and Steven L. Waslander. Categorical depth distribution network for monocular 3d object detection. CVPR, 2021.
  32. 32.Daniel Krajzewicz, Jakob Erdmann, Michael Behrisch, and Laura Bieker. Recent development and applications of sumo-simulation of urban mobility. International journal on advances in systems and measurements, 5(3&4), 2012.
  33. 33.Pengxiang Wu, Siheng Chen, and Dimitris N. Metaxas. Motionnet: Joint perception and motion prediction for autonomous driving based on bird’s eye view maps. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11382–11392, 2020.
  34. 34.Shital Shah, Debadeepta Dey, Chris Lovett, and Ashish Kapoor. Airsim: High-fidelity visual and physical simulation for autonomous vehicles. In Field and service robotics, pages 621–635. Springer, 2018.
  35. 35.Alex H. Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. Pointpillars: Fast encoders for object detection from point clouds. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12689–12697, 2019.
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  37. 37.Fisher Yu, Dequan Wang, Evan Shelhamer, and Trevor Darrell. Deep layer aggregation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2403–2412, 2018.
  38. 38.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable DETR: Deformable transformers for end-to-end object detection. ICLR, 2021.
  39. 39.Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. arXiv preprint arXiv:1903.11027, 2019.

Citation

MLA
Hu, Y., et al. “Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps”. arXiv, 2022, http://arxiv.org/abs/2209.12836v1.
APA
Hu, Y., Fang, S., Lei, Z., Zhong, Y., & Chen, S. (2022). Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps. arXiv. http://arxiv.org/abs/2209.12836v1
Chicago
Hu, Y., S. Fang, Z. Lei, Y. Zhong, and S. Chen. 2022. “Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps”. arXiv. http://arxiv.org/abs/2209.12836v1.
Harvard
Hu, Y. et al. (2022) “Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2209.12836v1.
Vancouver
1. Hu Y, Fang S, Lei Z, Zhong Y, Chen S (2022) Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps. arXiv

BibTeX

@article{hu2022where2comm,
  title = {Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps},
  author = {Hu, Yue and Fang, Shaoheng and Lei, Zixing and Zhong, Yiqi and Chen, Siheng},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2209.12836v1},
  eprint = {2209.12836}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors