Zero-Shot Object Counting

Jingyi XuHieu LeVu NguyenViresh RanjanDimitris Samaras

article2023CVPR81 citations

Introduces a zero-shot object counting framework that automatically localizes and selects optimal visual exemplars using only class names and a learned error predictor, eliminating the need for human-annotated bounding boxes at test time.

Listen

Automated computer vision systems increasingly require the ability to count specific objects across diverse, complex environments such as automated wildlife monitoring, inventory tracking, and anomaly detection. Traditional counting methods either specialize in a single pre-trained category or require human operators to manually provide visual bounding boxes—known as exemplars—for every new item to be counted. Existing automated alternatives only count the single most common object in an image and cannot be directed to count a specific category. This lack of flexible, fully automated counting creates operational bottlenecks and increases manual labor costs in autonomous workflows.

The article introduces and evaluates a practical new computer vision task called zero-shot object counting. The primary objective is to demonstrate a fully automated framework that accurately counts instances of any specified target category using only text input (the class name) without requiring human-annotated image examples.

To achieve this, the researchers developed a two-stage patch selection framework coupled with a base counting model. The system first generates a representative visual feature prototype from the text class name using a conditional generative model. It samples multiple image regions (candidate patches) and identifies those closest to the prototype in feature space. In the second step, an error-prediction model evaluates these candidates and selects the patches that will minimize counting errors. The article evaluated this approach on the FSC-147 benchmark dataset—comprising 6,135 images across 147 distinct categories—and compared performance against existing human-guided and automated counting methods.

The key findings demonstrate that this text-driven approach significantly outperforms previous exemplar-free counting baselines and nearly matches human-annotated performance. On the FSC-147 test benchmark, the proposed method reduced root mean squared error by 14.52 points compared to the existing exemplar-free alternative (RepRPN). Additionally, while other automated proposal methods caused severe performance drops when replacing human bounding boxes, the proposed patch selection method achieved counting accuracy comparable to human-provided inputs, showing only a minimal 1.41-point increase in mean absolute error. Ablation tests confirmed that both the prototype-based patch localization and the error-prediction selection are essential, with each component independently improving baseline accuracy by over 6 to 7 error points. Furthermore, the selected patches successfully generalized to enhance other existing vision counters (such as FamNet and BMNet) and demonstrated reliable counting in complex multi-class images where multiple distinct categories coexist.

These findings indicate that organizations can eliminate the operational overhead of human-in-the-loop annotations for visual counting tasks without sacrificing counting reliability. Autonomous systems can be deployed more rapidly into novel environments, reducing data preparation costs and enabling flexible, targeted tracking across diverse object types.

Decision-makers and engineering teams looking to adopt automated visual counting should consider integrating this text-directed patch selection pipeline into existing vision infrastructure to replace manual box-drawing workflows. For deployment in mission-critical applications, teams should conduct initial pilot evaluations in target operational domains to calibrate sensitivity thresholds, particularly in environments with extreme object occlusion or rare visual categories.

The study's results offer high confidence on natural images within the benchmark's distribution, though performance relies on the descriptive quality of underlying language-vision embeddings (such as CLIP) and pre-trained image representations. Caution is warranted when applying the system to highly specialized or abstract categories where text descriptions do not align cleanly with standard visual feature spaces.

Cover for Zero-Shot Object Counting

Citation

MLA
Xu, J., et al. “Zero-shot Object Counting”. arXiv, 2023, http://arxiv.org/abs/2303.02001v2.
APA
Xu, J., Le, H., Nguyen, V., Ranjan, V., & Samaras, D. (2023). Zero-shot Object Counting. arXiv. http://arxiv.org/abs/2303.02001v2
Chicago
Xu, J., H. Le, V. Nguyen, V. Ranjan, and D. Samaras. 2023. “Zero-shot Object Counting”. arXiv. http://arxiv.org/abs/2303.02001v2.
Harvard
Xu, J. et al. (2023) “Zero-shot Object Counting”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.02001v2.
Vancouver
1. Xu J, Le H, Nguyen V, Ranjan V, Samaras D (2023) Zero-shot Object Counting. arXiv

BibTeX

@article{xu2023zero,
  title = {Zero-shot Object Counting},
  author = {Xu, Jingyi and Le, Hieu and Nguyen, Vu and Ranjan, Viresh and Samaras, Dimitris},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.02001v2},
  eprint = {2303.02001}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE