Built independently by an author, for readers. Read the story and support ChapterPal

keyword

object proposals

Object proposals are candidate image regions, typically parameterized as bounding boxes or spatial segments, that are hypothesized to contain an object of interest rather than background. In computer vision and visual recognition pipelines, they serve to reduce computational complexity by narrowing the search space from an exhaustive sliding-window scan across every pixel coordinate and scale to a focused set of high-probability regions. These candidate regions are generated based on category-independent objectness cues using either low-level image features, such as edge contours and color groupings, or learned deep neural network modules like region proposal networks. Subsequent stages in an object detection or vision-language framework then score, refine, and classify these proposals into specific semantic categories.

28 items

ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension

ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension

Sanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner, Sameer Singh, Anna Rohrbach

Why you should read this

Presents ReCLIP, a zero-shot referring expression comprehension method that repurposes pre-trained vision-language models through isolated proposal scoring and explicit spatial relation resolution to outperform prior zero-shot baselines and rival out-of-domain supervised models.

Training a referring expression comprehension (ReC) model for a new visual domain requires collecting referring expressions, and potentially corresponding bounding boxes, for images in the domain. While large-scale pre-trained models are useful for image classification across domains, it remains unclear if they can be applied in a zero-shot manner to more complex tasks like ReC. We present ReCLIP, a simple but strong zero-shot baseline that repurposes CLIP, a state-of-the-art large-scale model, for ReC. Motivated by the close connection between ReC and CLIP’s contrastive pre-training objective, the first component of ReCLIP is a region-scoring method that isolates object proposals via cropping and blurring, and passes them to CLIP. However, through controlled experiments on a synthetic dataset, we find that CLIP is largely incapable of performing spatial reasoning off-the-shelf. Thus, the second component of ReCLIP is a spatial relation resolver that handles several types of spatial relations. We reduce the gap between zero-shot baselines from prior work and supervised models by as much as 29% on RefCOCOg, and on RefGTA (video game imagery), ReCLIP’s relative improvement over supervised ReC models trained on real images is 8%.

Added

2026-10-05

PROB: Probabilistic Objectness for Open World Object Detection

PROB: Probabilistic Objectness for Open World Object Detection

Orr Zohar, Kuan-Chieh Wang, Serena Yeung

OrganizationsStanford University

Why you should read this

Presents a probabilistic framework that models feature-space objectness distributions to distinguish unknown objects from background without pseudo-labels, doubling unknown object recall over prior open-world detection methods.

Open World Object Detection (OWOD) is a new and challenging computer vision task that bridges the gap between classic object detection (OD) benchmarks and object detection in the real world. In addition to detecting and classifying seen/labeled objects, OWOD algorithms are expected to detect novel/unknown objects - which can be classified and incrementally learned. In standard OD, object proposals not overlapping with a labeled object are automatically classified as background. Therefore, simply applying OD methods to OWOD fails as unknown objects would be predicted as background. The challenge of detecting unknown objects stems from the lack of supervision in distinguishing unknown objects and background object proposals. Previous OWOD methods have attempted to overcome this issue by generating supervision using pseudo-labeling - however, unknown object detection has remained low. Probabilistic/generative models may provide a solution for this challenge. Herein, we introduce a novel probabilistic framework for objectness estimation, where we alternate between probability distribution estimation and objectness likelihood maximization of known objects in the embedded feature space - ultimately allowing us to estimate the objectness probability of different proposals. The resulting Probabilistic Objectness transformer-based open-world detector, PROB, integrates our framework into traditional object detection models, adapting them for the open-world setting. Comprehensive experiments on OWOD benchmarks show that PROB outperforms all existing OWOD methods in both unknown object detection (~ 2× unknown recall) and known object detection (~ 10% mAP). Our code is available at https://github.com/orrzohar/PROB.

Added

2026-09-26

The Secrets of Salient Object Segmentation

The Secrets of Salient Object Segmentation

Yin Li, Xiaodi Hou, Christof Koch, James M. Rehg, Alan L. Yuille

OrganizationsAllen Institute for Brain ScienceCalifornia Institute of TechnologyGeorgia Institute of TechnologyUniversity of California, Los Angeles

Why you should read this

Exposes critical design biases in salient object benchmarks and introduces a unified dataset with joint fixation and segmentation ground truth alongside a method that directly connects human visual attention to object segmentation.

In this paper we provide an extensive evaluation of fixation prediction and salient object segmentation algorithms as well as statistics of major datasets. Our analysis identifies serious design flaws of existing salient object benchmarks, called the dataset design bias, by over emphasizing the stereotypical concepts of saliency. The dataset design bias does not only create the discomforting disconnection between fixations and salient object segmentation, but also misleads the algorithm designing. Based on our analysis, we propose a new high quality dataset that offers both fixation and salient object segmentation ground-truth. With fixations and salient object being presented simultaneously, we are able to bridge the gap between fixations and salient objects, and propose a novel method for salient object segmentation. Finally, we report significant benchmark progress on three existing datasets of segmenting salient objects

Added

2026-09-25

From captions to visual concepts and back

From captions to visual concepts and back

Hao Fang, Saurabh Gupta, F. Iandola, R. Srivastava, L. Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C. Platt, C. L. Zitnick, G. Zweig

OrganizationsDalle Molle Institute for Artificial Intelligence ResearchGoogleMetaMicrosoftUniversity of California BerkeleyUniversity of Washington

Why you should read this

Presents an image captioning pipeline that learns visual concept detectors directly from weakly-supervised captions via multiple instance learning, generates candidate descriptions with a maximum-entropy language model, and selects the best description using a deep multimodal similarity model.

This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in captions, including many different parts of speech such as nouns, verbs, and adjectives. The word detector outputs serve as conditional inputs to a maximum-entropy language model. The language model learns from a set of over 400,000 image descriptions to capture the statistics of word usage. We capture global semantics by re-ranking caption candidates using sentence-level features and a deep multimodal similarity model. Our system is state-of-the-art on the official Microsoft COCO benchmark, producing a BLEU-4 score of 29.1%. When human judges compare the system captions to ones written by other people on our held-out test set, the system captions have equal or better quality 34% of the time.

Added

2026-09-25

Single-Shot Refinement Neural Network for Object Detection

Single-Shot Refinement Neural Network for Object Detection

Shifeng Zhang, Longyin Wen, Xiao Bian, Zhen Lei, Stan Z. Li

OrganizationsGeneral Electric CompanyInstitute of Automation, Chinese Academy of SciencesUniversity of Chinese Academy of Sciences

Why you should read this

Proposes RefineDet, an object detector that combines the high accuracy of two-stage methods with the fast inference of single-stage models by using anchor refinement and feature transfer modules to filter false positives and optimize bounding boxes before final classification.

For object detection, the two-stage approach (e.g., Faster R-CNN) has been achieving the highest accuracy, whereas the one-stage approach (e.g., SSD) has the advantage of high efficiency. To inherit the merits of both while overcoming their disadvantages, in this paper, we propose a novel single-shot based detector, called RefineDet, that achieves better accuracy than two-stage methods and maintains comparable efficiency of one-stage methods. RefineDet consists of two inter-connected modules, namely, the anchor refinement module and the object detection module. Specifically, the former aims to (1) filter out negative anchors to reduce search space for the classifier, and (2) coarsely adjust the locations and sizes of anchors to provide better initialization for the subsequent regressor. The latter module takes the refined anchors as the input from the former to further improve the regression and predict multi-class label. Meanwhile, we design a transfer connection block to transfer the features in the anchor refinement module to predict locations, sizes and class labels of objects in the object detection module. The multi-task loss function enables us to train the whole network in an end-to-end way. Extensive experiments on PASCAL VOC 2007, PASCAL VOC 2012, and MS COCO demonstrate that RefineDet achieves state-of-the-art detection accuracy with high efficiency. Code is available at this https URL

Added

2026-09-25

Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin Cui

OrganizationsGoogleNVIDIA

Why you should read this

Introduces ViLD, a distillation framework that transfers multimodal knowledge from pretrained vision-language models into two-stage object detectors, enabling the detection of novel categories specified by arbitrary text without requiring additional bounding-box annotations.

We aim at advancing open-vocabulary object detection, which detects objects described by arbitrary text inputs. The fundamental challenge is the availability of training data. It is costly to further scale up the number of classes contained in existing object detection datasets. To overcome this challenge, we propose ViLD, a training method via Vision and Language knowledge Distillation. Our method distills the knowledge from a pretrained open-vocabulary image classification model (teacher) into a two-stage detector (student). Specifically, we use the teacher model to encode category texts and image regions of object proposals. Then we train a student detector, whose region embeddings of detected boxes are aligned with the text and image embeddings inferred by the teacher. We benchmark on LVIS by holding out all rare categories as novel categories that are not seen during training. ViLD obtains 16.1 mask APr_r with a ResNet-50 backbone, even outperforming the supervised counterpart by 3.8. When trained with a stronger teacher model ALIGN, ViLD achieves 26.3 APr_r. The model can directly transfer to other datasets without finetuning, achieving 72.2 AP50_{50} on PASCAL VOC, 36.6 AP on COCO and 11.8 AP on Objects365. On COCO, ViLD outperforms the previous state-of-the-art by 4.8 on novel AP and 11.4 on overall AP. Code and demo are open-sourced at this https URL.

Added

2026-09-25

Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, Marcus Rohrbach

OrganizationsMax Planck Institute for InformaticsSony CorporationUniversity of California Berkeley

Why you should read this

Introduces Multimodal Compact Bilinear pooling to efficiently capture high-dimensional cross-modal interactions, achieving state-of-the-art performance in visual question answering and visual grounding.

Modeling textual or visual information with vector representations trained from large language or visual datasets has been successfully explored in recent years. However, tasks such as visual question answering require combining these vector representations with each other. Approaches to multimodal pooling include element-wise product or sum, as well as concatenation of the visual and textual representations. We hypothesize that these methods are not as expressive as an outer product of the visual and textual vectors. As the outer product is typically infeasible due to its high dimensionality, we instead propose utilizing Multimodal Compact Bilinear pooling (MCB) to efficiently and expressively combine multimodal features. We extensively evaluate MCB on the visual question answering and grounding tasks. We consistently show the benefit of MCB over ablations without MCB. For visual question answering, we present an architecture which uses MCB twice, once for predicting attention over spatial features and again to combine the attended representation with the question representation. This model outperforms the state-of-the-art on the Visual7W dataset and the VQA challenge.

Added

2026-09-25

Creative Commons License
Sparse R-CNN: End-to-End Object Detection with Learnable Proposals

Sparse R-CNN: End-to-End Object Detection with Learnable Proposals

Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, Ping Luo

OrganizationsByteDanceTongji UniversityUniversity of California BerkeleyUniversity of Hong Kong

Why you should read this

Introduces Sparse R-CNN, a purely sparse object detection framework that replaces dense anchor boxes with a small set of learnable proposals, eliminating hand-designed candidates and non-maximum suppression while achieving competitive accuracy and speed on the COCO benchmark.

We present Sparse R-CNN, a purely sparse method for object detection in images. Existing works on object detection heavily rely on dense object candidates, such as kk anchor boxes pre-defined on all grids of image feature map of size H×WH\times W. In our method, however, a fixed sparse set of learned object proposals, total length of NN, are provided to object recognition head to perform classification and location. By eliminating HWkHWk (up to hundreds of thousands) hand-designed object candidates to NN (e.g. 100) learnable proposals, Sparse R-CNN completely avoids all efforts related to object candidates design and many-to-one label assignment. More importantly, final predictions are directly output without non-maximum suppression post-procedure. Sparse R-CNN demonstrates accuracy, run-time and training convergence performance on par with the well-established detector baselines on the challenging COCO dataset, e.g., achieving 45.0 AP in standard 3×3\times training schedule and running at 22 fps using ResNet-50 FPN model. We hope our work could inspire re-thinking the convention of dense prior in object detectors. The code is available at: this https URL.

Added

2026-09-24

UnitBox: An Advanced Object Detection Network

UnitBox: An Advanced Object Detection Network

Jiahui Yu, Yuning Jiang, Zhangyang Wang, Zhimin Cao, Thomas S. Huang

OrganizationsMegvii TechnologyUniversity of Illinois Urbana-Champaign

Why you should read this

Proposes the Intersection over Union (IoU) loss function to optimize bounding box coordinates as an integrated unit rather than independent variables, overcoming traditional regression limitations to achieve superior localization accuracy in object detection.

In present object detection systems, the deep convolutional neural networks (CNNs) are utilized to predict bounding boxes of object candidates, and have gained performance advantages over the traditional region proposal methods. However, existing deep CNN methods assume the object bounds to be four independent variables, which could be regressed by the ℓ2\ell_2 loss separately. Such an oversimplified assumption is contrary to the well-received observation, that those variables are correlated, resulting to less accurate localization. To address the issue, we firstly introduce a novel Intersection over Union (IoUIoU) loss function for bounding box prediction, which regresses the four bounds of a predicted box as a whole unit. By taking the advantages of IoUIoU loss and deep fully convolutional networks, the UnitBox is introduced, which performs accurate and efficient localization, shows robust to objects of varied shapes and scales, and converges fast. We apply UnitBox on face detection task and achieve the best performance among all published methods on the FDDB benchmark.

Added

2026-09-24

Salient Object Detection: A Benchmark

Salient Object Detection: A Benchmark

Ali Borji, Ming-Ming Cheng, Huaizu Jiang, Jia Li

OrganizationsBeihang UniversityNankai UniversityUniversity of Wisconsin-MilwaukeeXi'an Jiaotong University

Why you should read this

Establishes a comprehensive benchmark by evaluating forty models across six datasets, analyzing factors like center bias and scene complexity to expose key failure modes and guide future salient object detection research.

We extensively compare, qualitatively and quantitatively, 40 state-of-the-art models (28 salient object detection, 10 fixation prediction, 1 objectness, and 1 baseline) over 6 challenging datasets for the purpose of benchmarking salient object detection and segmentation methods. From the results obtained so far, our evaluation shows a consistent rapid progress over the last few years in terms of both accuracy and running time. The top contenders in this benchmark significantly outperform the models identified as the best in the previous benchmark conducted just two years ago. We find that the models designed specifically for salient object detection generally work better than models in closely related areas, which in turn provides a precise definition and suggests an appropriate treatment of this problem that distinguishes it from other problems. In particular, we analyze the influences of center bias and scene complexity in model performance, which, along with the hard cases for state-of-the-art models, provide useful hints towards constructing more challenging large scale datasets and better saliency models. Finally, we propose probable solutions for tackling several open problems such as evaluation scores and dataset bias, which also suggest future research directions in the rapidly-growing field of salient object detection.

Added

2026-09-18

A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation

A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation

Federico Perazzi, J. Pont-Tuset, B. McWilliams, L. Gool, M. Gross, A. Sorkine-Hornung

OrganizationsDisney ResearchETH Zurich

Why you should read this

Introduces the densely annotated DAVIS video-segmentation benchmark and complementary spatial, contour, and temporal metrics that expose the strengths and weaknesses of current methods.

Over the years, datasets and benchmarks have proven their fundamental importance in computer vision research, enabling targeted progress and objective comparisons in many fields. At the same time, legacy datasets may impend the evolution of a field due to saturated algorithm performance and the lack of contemporary, high quality data. In this work we present a new benchmark dataset and evaluation methodology for the area of video object segmentation. The dataset, named DAVIS (Densely Annotated VIdeo Segmentation), consists of fifty high quality, Full HD video sequences, spanning multiple occurrences of common video object segmentation challenges such as occlusions, motion-blur and appearance changes. Each video is accompanied by densely annotated, pixel-accurate and per-frame ground truth segmentation. In addition, we provide a comprehensive analysis of several state-of-the-art segmentation approaches using three complementary metrics that measure the spatial extent of the segmentation, the accuracy of the silhouette contours and the temporal coherence. The results uncover strengths and weaknesses of current approaches, opening up promising directions for future works.

Added

2026-09-14

Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, Svetlana Lazebnik

OrganizationsFundación Universitaria Konrad LorenzUniversity of Illinois Urbana-Champaign

Why you should read this

Introduces the Flickr30k Entities benchmark by grounding over 240,000 caption phrases to image bounding boxes, establishing a standard dataset and baseline for phrase localization and vision-language grounding.

The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains, linking mentions of the same entities across different captions for the same image, and associating them with 276k manually annotated bounding boxes. Such annotations are essential for continued progress in automatic image description and grounded language understanding. They enable us to define a new benchmark for localization of textual entity mentions in an image. We present a strong baseline for this task that combines an image-text embedding, detectors for common objects, a color classifier, and a bias towards selecting larger objects. While our baseline rivals in accuracy more complex state-of-the-art models, we show that its gains cannot be easily parlayed into improvements on such tasks as image-sentence retrieval, thus underlining the limitations of current methods and the need for further research.

Added

2026-09-14

Training Region-Based Object Detectors with Online Hard Example Mining

Training Region-Based Object Detectors with Online Hard Example Mining

Abhinav Shrivastava, Abhinav Gupta, Ross Girshick

OrganizationsCarnegie Mellon UniversityMeta

Why you should read this

Proposes Online Hard Example Mining (OHEM) to automatically select difficult region proposals during ConvNet training, eliminating heuristic sampling hyperparameters while boosting object detection accuracy across standard benchmarks.

The field of object detection has made significant advances riding on the wave of region-based ConvNets, but their training procedure still includes many heuristics and hyperparameters that are costly to tune. We present a simple yet surprisingly effective online hard example mining (OHEM) algorithm for training region-based ConvNet detectors. Our motivation is the same as it has always been -- detection datasets contain an overwhelming number of easy examples and a small number of hard examples. Automatic selection of these hard examples can make training more effective and efficient. OHEM is a simple and intuitive algorithm that eliminates several heuristics and hyperparameters in common use. But more importantly, it yields consistent and significant boosts in detection performance on benchmarks like PASCAL VOC 2007 and 2012. Its effectiveness increases as datasets become larger and more difficult, as demonstrated by the results on the MS COCO dataset. Moreover, combined with complementary advances in the field, OHEM leads to state-of-the-art results of 78.9% and 76.3% mAP on PASCAL VOC 2007 and 2012 respectively.

Added

2026-09-14

Deep Learning for Generic Object Detection: A Survey

Deep Learning for Generic Object Detection: A Survey

Li Liu, Wanli Ouyang, Xiaogang Wang, Paul Fieguth, Jie Chen, Xinwang Liu, Matti Pietikäinen

OrganizationsNational University of Defense TechnologyThe Chinese University of Hong KongUniversity of OuluUniversity of SydneyUniversity of Waterloo

Why you should read this

Systematizes over three hundred deep learning studies on generic object detection, providing a structured comparison of detector frameworks, proposal generation, context modeling, and training strategies to guide computer vision researchers.

Object detection, one of the most fundamental and challenging problems in computer vision, seeks to locate object instances from a large number of predefined categories in natural images. Deep learning techniques have emerged as a powerful strategy for learning feature representations directly from data and have led to remarkable breakthroughs in the field of generic object detection. Given this period of rapid evolution, the goal of this paper is to provide a comprehensive survey of the recent achievements in this field brought about by deep learning techniques. More than 300 research contributions are included in this survey, covering many aspects of generic object detection: detection frameworks, object feature representation, object proposal generation, context modeling, training strategies, and evaluation metrics. We finish the survey by identifying promising directions for future research.

Added

2026-09-14

Edge Boxes: Locating Object Proposals from Edges

Edge Boxes: Locating Object Proposals from Edges

C. Lawrence Zitnick, Piotr Dollár

OrganizationsMicrosoft

Why you should read this

Proposes an efficient object proposal method that scores candidate bounding boxes by measuring wholly enclosed edge contours while penalizing boundary-straddling edges, achieving state-of-the-art detection recall in a fraction of a second.

The use of object proposals is an effective recent approach for increasing the computational efficiency of object detection. We propose a novel method for generating object bounding box proposals using edges. Edges provide a sparse yet informative representation of an image. Our main observation is that the number of contours that are wholly contained in a bounding box is indicative of the likelihood of the box containing an object. We propose a simple box objectness score that measures the number of edges that exist in the box minus those that are members of contours that overlap the box's boundary. Using efficient data structures, millions of candidate boxes can be evaluated in a fraction of a second, returning a ranked set of a few thousand top-scoring proposals. Using standard metrics, we show results that are significantly more accurate than the current state-of-the-art while being faster to compute. In particular, given just 1000 proposals we achieve over 96% object recall at overlap threshold of 0.5 and over 75% recall at the more challenging overlap of 0.7. Our approach runs in 0.25 seconds and we additionally demonstrate a near real-time variant with only minor loss in accuracy.

Added

2026-09-14