Foreground localization is a computer vision process that involves identifying and determining the spatial extent or pixel-level regions of target objects and salient visual elements within an image while distinguishing them from the surrounding background. In weakly supervised learning paradigms, such as semantic segmentation and object detection where detailed pixel-level annotations are unavailable, foreground localization commonly relies on activation maps and feature representations derived from image-level category labels to generate initial spatial cues. A key objective in foreground localization is expanding beyond the most discriminative parts of an object to capture its entire semantic structure, providing reliable pseudo-ground-truth masks for downstream dense prediction models.