keyword
CoDet
CoDet is a computer vision framework designed for open-vocabulary object detection that aligns image regions with textual words by discovering co-occurring objects across multiple image-text pairs. Rather than relying strictly on pre-aligned vision-language embedding spaces, the approach groups images that share common concepts in their captions and uses visual similarity to identify the recurring visual regions corresponding to those words. By framing region-word alignment as a co-occurring object discovery problem, CoDet enables detectors to learn fine-grained visual representations from large-scale image-text data without requiring extensive bounding-box annotations, improving both localization accuracy and the model ability to recognize novel object categories.
1 item

