Structure-Measure: A New Way to Evaluate Foreground Maps
Deng-Ping FanMing-Ming ChengYun LiuTao LiAli Borji
Proposes the Structure-measure, an evaluation metric that combines region-aware and object-aware structural similarity to assess salient object detection algorithms more accurately than conventional pixel-wise measures.
Accurately evaluating foreground and salient object detection models is vital across modern computer vision tasks, including semantic segmentation, image retrieval, and object detection. However, standard evaluation metrics rely on pixel-level error comparisons and fail to capture global object structure. This structural blind spot causes conventional metrics to misrank predicted maps, often scoring degraded or incomplete object shapes higher than well-preserved detections.
The article aims to design and validate a novel evaluation metric, termed Structure-measure, that assesses structural similarity between non-binary predicted maps and ground-truth human annotations. The proposed approach balances both region-aware structural information, which recursively evaluates sub-regions, and object-aware similarity, which evaluates global foreground-background contrast and uniform saliency distribution.
To establish credibility, the authors validated the metric across five benchmark image datasets using ten state-of-the-art salient object detection models and five diagnostic meta-measures. In these tests, Structure-measure consistently outperformed existing metrics, demonstrating superior ranking consistency with practical applications and reducing incorrect ground-truth match errors by 18% to 69% compared to the previous leading metric. Furthermore, in a controlled behavioral study of 45 human participants across 50 paired evaluations, viewers agreed with Structure-measure rankings 64% to 74% of the time over traditional alternatives, all while maintaining high computational efficiency at roughly 5.3 milliseconds per image.
These findings indicate that relying on legacy pixel-based metrics creates significant risk of selecting suboptimal computer vision models in real-world deployments. Adopting Structure-measure provides engineering and development teams with a more reliable benchmark to guide model selection and improve downstream system performance without incurring high computational overhead.
The source recommends that practitioners and researchers adopt Structure-measure to evaluate and compare salient object detection models. Future efforts should focus on integrating this metric into training loss pipelines and expanding evaluations to broader segmentation tasks. Because the current formulation assumes binary ground-truth reference maps and relies on specific parameter weightings, teams should ensure these operational assumptions align with their specific production environments.
- Paper: Global contrast based salient region detection, Ming-Ming Cheng et al. (2011). This paper establishes foundational global contrast formulations and benchmark methodologies for salient object detection and foreground map generation.
- Paper: A Model of Saliency-Based Visual Attention for Rapid Scene Analysis, L. Itti et al. (1998). It introduces the classical bottom-up visual saliency framework, providing the essential conceptual grounding for visual attention and saliency map formation.
- Paper: Graph-Based Visual Saliency, Jonathan Harel et al. (2006). It provides standard graph-based saliency modeling and evaluation conventions against human fixation ground truth that underpin saliency research.
- Paper: Contour Detection and Hierarchical Image Segmentation, Pablo Arbeláez et al. (2011). It outlines fundamental region- and contour-based segmentation evaluation metrics, motivating the need to measure structural and boundary similarities.
- Paper: A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation, Federico Perazzi et al. (2016). It details standard evaluation methodologies and boundary-versus-region similarity metrics for foreground object segmentation.
- Paper: U2-Net: Going Deeper with Nested U-Structure for Salient Object Detection, Xuebin Qin et al. (2020). It designs a deep nested U-structure network for salient object detection and directly utilizes the Structure-measure (S-measure) to benchmark foreground map quality.
- Paper: The Unreasonable Effectiveness of Deep Features as a Perceptual Metric, Richard Zhang et al. (2018). It advances the broader study of human perceptual similarity by demonstrating how deep neural representations align with visual structure better than traditional pixel-wise metrics.
