Optimal Transport Minimization: Crowd Localization on Density Maps for Semi-Supervised Counting
Wei LinAntoni B. Chan
Proposes a parameter-free optimal transport minimization algorithm that converts predicted crowd density maps into precise point locations by minimizing Sinkhorn distance, enabling hard pseudo-label generation and confidence-weighted training for semi-supervised crowd counting.
Crowd analysis in computer vision is critical for public safety, surveillance, and traffic management. While modern deep neural networks effectively estimate overall crowd numbers through blurry density maps, they struggle to pinpoint the precise physical locations of individual people. Consequently, organizations often maintain separate, redundant computer vision systems for counting and localization. Additionally, training high-performing crowd models requires large, densely labeled datasets, which are expensive and labor-intensive to annotate.
The article develops and evaluates a parameter-free method called Optimal Transport Minimization (OT-M) to extract exact individual coordinates directly from crowd density maps without additional training. Furthermore, it demonstrates how this method can generate hard point annotations on unlabeled images to improve semi-supervised crowd counting.
The researchers designed an iterative algorithm that minimizes the mathematical transport cost between a continuous density map and discrete target points. Using this algorithm, they built a semi-supervised learning framework where a teacher network generates point locations for unlabeled images to train a student network. To mitigate incorrect pseudo-annotations, they introduced a confidence weighting mechanism. The framework was evaluated across multiple standard public crowd benchmarks, including UCF-QNRF, NWPU-Crowd, ShanghaiTech, and JHU++, across settings where only 5%, 10%, or 40% of training images had ground-truth labels.
The evaluation revealed several key findings. First, OT-M significantly outperformed existing density-based localization techniques, achieving an F-measure of 0.912 on UCF-QNRF compared to 0.840 for Gaussian mixture models and 0.807 for local peak detection. Second, in semi-supervised counting with limited labeled data (5% and 10%), the OT-M framework achieved the lowest counting errors and highest stability across all tested datasets. Third, ablation experiments showed that training with exact point annotations combined with confidence weighting reduced counting error metrics by over 12% compared to standard soft density supervision and unweighted losses.
These results demonstrate that organizations do not need separate, specialized networks to count and locate individuals simultaneously, lowering computational overhead and deployment costs. The approach also substantially reduces manual data annotation expenses, as models can be trained effectively on datasets where 90% to 95% of the images lack human annotations. While the algorithm's output quality remains constrained by the resolution and precision of the underlying density maps, the mathematical convergence and consistent empirical performance provide high confidence in adopting this method for operational crowd monitoring and semi-supervised computer vision pipelines.
- Paper: Computational Optimal Transport, Gabriel Peyré et al. (2018). It provides the core mathematical theory and computational algorithms for optimal transport, establishing the foundational principles that the source adapts for continuous-to-discrete density map minimization.
- Paper: Learning To Count Objects in Images, V. Lempitsky et al. (2010). It introduces the foundational paradigm of converting point/dot annotations into continuous spatial density maps for visual counting.
- Paper: Single-Image Crowd Counting via Multi-Column Convolutional Neural Network, Yingying Zhang et al. (2016). It establishes standard deep convolutional density map estimation and provides benchmark datasets (like ShanghaiTech) utilized directly in the source paper.
- Paper: CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes, Yuhong Li et al. (2018). It presents widely adopted dilated convolutional neural network architectures for generating high-resolution crowd density maps evaluated by the source.
- Paper: On the Efficiency of Entropic Regularized Algorithms for Optimal Transport, Tianyi Lin et al. (2022). It develops fast entropic regularized algorithms for discrete optimal transport, offering essential background on efficient iterative transport cost optimization.
- Paper: Zero-Shot Object Counting, Jingyi Xu et al. (2023). It extends object counting beyond supervised and semi-supervised crowd scenarios into a zero-shot, text-prompted paradigm across diverse open-world categories.
