Single Domain Generalization for Crowd Counting

Zhuoxuan PengS.-H. Gary Chan

article2024CVPR55 citations

Develops MPCount, a single-domain generalization framework that overcomes domain shift and label ambiguity in crowd counting by coupling memory-based invariant feature reconstruction with patch-wise classification to accurately estimate densities across unseen environments.

Listen

Automated crowd counting using computer vision is essential for public safety, urban management, and operational planning. Standard systems rely on deep learning models trained to estimate crowd density maps from images. However, these systems frequently fail when deployed in real-world environments because variations in camera angles, scene structures, and weather conditions cause severe performance drops. While existing solutions require either target-environment data for fine-tuning or broad and diverse training datasets, real-world deployments rarely have access to target data beforehand and often only possess limited, narrowly distributed training imagery.

The article introduces and evaluates MPCount, a novel artificial intelligence framework designed for single-domain generalization in crowd counting. The primary objective is to demonstrate that an automated counting model can be trained on a single, narrowly defined environment (such as snowy or street scenes) and still accurately estimate crowd sizes in entirely unobserved, different environments without requiring target-site adaptation.

To achieve this, the approach introduces an attention-based memory bank that stores domain-invariant visual representations and reconstructs continuous crowd density values. It integrates a content error mask to remove style discrepancies caused by environmental shifts and an attention consistency loss to maintain uniform feature learning. Additionally, MPCount tackles label ambiguity—where head annotations and background pixels look similar across scenes—by adding an auxiliary patch-wise classification task that divides images into coarse grid patches to reliably distinguish crowd presence from empty background before final density estimation. The framework was evaluated across standard benchmark datasets (ShanghaiTech and UCF-QNRF) and challenging, narrowly distributed conditions within the JHU-Crowd++ dataset.

The empirical findings demonstrate significant improvements over existing methods. When trained on narrow single domains, MPCount reduced counting errors against leading benchmark techniques by 21.8% when transferring from snow to fog/haze scenarios and by 18.6% when transferring from fog/haze to snow. On standard cross-dataset benchmarks, it reduced counting errors by 18.2% when generalizing from street views to dense web imagery and by 9.5% in the reverse direction. Furthermore, MPCount consistently matched or surpassed specialized domain adaptation methods that had direct access to target domain images during training. Ablation studies confirmed that combining the memory module with patch-wise classification produced the most substantial performance gains.

These results indicate that organizations can successfully deploy vision-based crowd counting tools across diverse, unpredictable environments without collecting target-site training images or performing ongoing model fine-tuning. This significantly lowers operational costs, eliminates data collection risks, and accelerates deployment timelines in security and venue management operations. The findings also demonstrate that framing background filtering as a coarse patch classification task is far more effective at resolving label ambiguity than standard pixel-level regression.

Organizations seeking to implement zero-adaptation crowd monitoring systems should adopt patch-filtered memory architectures like MPCount. Next operational steps should involve conducting field pilot tests in operational surveillance networks and evaluating computational performance across edge hardware. While the reported experimental gains provide high confidence in the model's architectural advantages, decision-makers should note that evaluations were conducted on static 2D image benchmarks, and further testing is warranted for continuous video feeds and extreme lighting shifts.

No sufficiently relevant recommendations were found.

Cover for Single Domain Generalization for Crowd Counting

Abstract

Due to its promising results, density map regression has been widely employed for image-based crowd counting. The approach, however, often suffers from severe performance degradation when tested on data from unseen scenarios, the so-called "domain shift" problem. To address the problem, we investigate in this work single domain generalization (SDG) for crowd counting. The existing SDG approaches are mainly for image classification and segmentation, and can hardly be extended to our case due to its regression nature and label ambiguity (i.e., ambiguous pixel-level ground truths). We propose MPCount, a novel effective SDG approach even for narrow source distribution. MPCount stores diverse density values for density map regression and reconstructs domain-invariant features by means of only one memory bank, a content error mask and attention consistency loss. By partitioning the image into grids, it employs patch-wise classification as an auxiliary task to mitigate label ambiguity. Through extensive experiments on different datasets, MPCount is shown to significantly improve counting accuracy compared to the state of the art under diverse scenarios unobserved in the training data characterized by narrow source distribution. Code is available at this https URL.

Citation

MLA
Peng, Z., and S.-H. G. Chan. “Single Domain Generalization for Crowd Counting”. arXiv, 2024, http://arxiv.org/abs/2403.09124v2.
APA
Peng, Z., & Chan, S.-H. G. (2024). Single Domain Generalization for Crowd Counting. arXiv. http://arxiv.org/abs/2403.09124v2
Chicago
Peng, Z., and S.-H. G. Chan. 2024. “Single Domain Generalization for Crowd Counting”. arXiv. http://arxiv.org/abs/2403.09124v2.
Harvard
Peng, Z. and Chan, S.-H.G. (2024) “Single Domain Generalization for Crowd Counting”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.09124v2.
Vancouver
1. Peng Z, Chan S-HG (2024) Single Domain Generalization for Crowd Counting. arXiv

BibTeX

@article{peng2024single,
  title = {Single Domain Generalization for Crowd Counting},
  author = {Peng, Zhuoxuan and Chan, S. -H. Gary},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.09124v2},
  eprint = {2403.09124}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/