Bridging the Gap between Classification and Localization for Weakly Supervised Object Localization

Eunji KimSiwon KimJungbeom LeeHyunwoo KimSungroh Yoon

article2022CVPR52 citations

Reveals that misaligned feature and classifier weight directions cause class activation maps to miss non-discriminative object parts, and proposes a feature direction alignment method alongside attentive dropout to achieve state-of-the-art weakly supervised object localization on CUB-200-2011 and ImageNet-1K.

Listen

Training computer vision models to locate objects within images typically requires manually drawing bounding boxes around every target, a process that is labor-intensive and costly. Weakly supervised object localization addresses this challenge by training models using only simple image-level category labels. However, existing standard approaches using class activation maps suffer from a fundamental drawback: they typically identify only the most prominent, discriminative parts of an object (such as the body or head of a bird) rather than outlining the entire object area (such as wings and legs).

The main objective of the article is to diagnose why class activation maps fail to capture full object extents and to evaluate a new training framework designed to align intermediate visual features with class classifiers. The article demonstrates that mathematically aligning spatial feature directions with class-specific weights substantially closes the performance gap between image classification and complete object localization.

To investigate and resolve this issue, the authors decomposed the standard activation map into two factors: the magnitude of regional feature activations and their directional alignment (cosine similarity) with class classifier weights. Based on this decomposition, they developed a single-model training method incorporating two core techniques: feature direction alignment loss, which forces target regions to align with class weights while suppressing background noise, and attentive dropout consistency, which stochastically removes peak activations to evenly spread feature emphasis across the target. Credibility was established through extensive benchmarking on two standard datasets—the fine-grained CUB-200-2011 dataset and the large-scale ImageNet-1K dataset—evaluated across common backbone architectures such as VGG16 and ResNet50.

The experimental findings show significant performance improvements across all benchmarks. On CUB-200-2011, the method achieved a top-1 localization accuracy of 70.83% with VGG16 and 73.16% with ResNet50, outperforming prior single-branch CAM-based state-of-the-art methods by 11.87 and over 13 percentage points, respectively. The approach also exceeded the performance of complex multi-branch architectures while using fewer computational resources. Under strict bounding box overlap thresholds (MaxBoxAccV2 at 0.7 IoU), the framework improved accuracy by 17.4 to 21.0 percentage points on CUB-200-2011. On the ImageNet-1K dataset, the method attained state-of-the-art results across most metrics, reaching 49.94% top-1 localization accuracy on VGG16 and 69.89% ground-truth localization accuracy on ResNet50.

These findings imply that organizations can achieve highly accurate object localization at substantially reduced data annotation costs, bypassing the need for manual bounding box labeling. Because the technique operates within standard single-network pipelines without adding secondary models or inference overhead, it avoids the latency and memory penalties associated with multi-branch systems. Practitioners can deploy more reliable visual recognition models while controlling development costs and computational budgets.

Engineering teams building vision systems should consider adopting feature direction alignment and attentive dropout mechanisms during the training of weakly supervised models. When implementing the method, practitioners should follow a staged training schedule—beginning with a warm-up phase focused on classification and dropout before activating directional alignment losses. Further work and pilot validations are suggested to refine the automated selection of balancing hyperparameters across varied industrial datasets.

The primary limitation of the proposed approach is the introduction of several balancing hyperparameters that govern loss weights and dropout thresholds. While the article demonstrates that sensitivity around key thresholds is relatively robust, extreme values can degrade localization quality. Confidence in the reported results is high, given the consistent state-of-the-art gains across diverse standard vision benchmarks and network architectures.

arXiv: 2204.00220
Cover for Bridging the Gap between Classification and Localization for Weakly Supervised Object Localization

Abstract

Weakly supervised object localization aims to find a target object region in a given image with only weak supervision, such as image-level labels. Most existing methods use a class activation map (CAM) to generate a localization map; however, a CAM identifies only the most discriminative parts of a target object rather than the entire object region. In this work, we find the gap between classification and localization in terms of the misalignment of the directions between an input feature and a class-specific weight. We demonstrate that the misalignment suppresses the activation of CAM in areas that are less discriminative but belong to the target object. To bridge the gap, we propose a method to align feature directions with a class-specific weight. The proposed method achieves a state-of-the-art localization performance on the CUB-200-2011 and ImageNet-1K benchmarks.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Finding the Gap with CAM Decomposition
  • 4. Bridging the Gap through Alignment
  • 4.1. Alignment of Feature Directions
  • 4.2. Consistency with Attentive Dropout
  • 4.3. Training Scheme
  • 5. Experiments
  • 5.1. Experimental Settings
  • 5.2. Comparison with State-of-the-art Methods
  • 5.3. Discussion
  • 5.4. Ablation Study
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Decomposition of Class Activation Maps into Feature Norm and Cosine Similarity

    definition

    Given an input image xx, a convolutional feature map F(x)∈RH×W×DF(x) \in \mathbb{R}^{H \times W \times D} prior to the global average pooling (GAP) layer, and the classification weight vector wc∈RDw_c \in \mathbb{R}^D of a linear classification layer for target class cc, the class activation map (CAM) value at spatial location index u∈{1,…,HW}u \in \{1, \dots, HW\} is defined by the dot product:

    CAMu(x)=wc⋅Fu(x)=∥wc∥∥Fu(x)∥wc⋅Fu(x)∥wc∥∥Fu(x)∥=∥wc∥FuSu\text{CAM}_u(x) = w_c \cdot F_u(x) = \|w_c\| \|F_u(x)\| \frac{w_c \cdot F_u(x)}{\|w_c\| \|F_u(x)\|} = \|w_c\| \mathcal{F}_u \mathcal{S}_u

    where:

    • Fu=∥Fu(x)∥\mathcal{F}_u = \|F_u(x)\| is the Euclidean norm of the spatial feature vector at position uu, forming a spatial norm map F∈RH×W\mathcal{F} \in \mathbb{R}^{H \times W}.
    • Su=S(wc,Fu(x))=wc⋅Fu(x)∥wc∥∥Fu(x)∥\mathcal{S}_u = S(w_c, F_u(x)) = \frac{w_c \cdot F_u(x)}{\|w_c\| \|F_u(x)\|} is the cosine similarity between the class weight vector wcw_c and the local feature vector Fu(x)F_u(x), forming a similarity map S∈RH×W\mathcal{S} \in \mathbb{R}^{H \times W}.
    • ∥wc∥\|w_c\| is the norm of the class weight vector, which is spatially invariant for a given class cc.

    In matrix notation across all spatial positions, CAM(x)=∥wc∥⋅(F⊙S)\text{CAM}(x) = \|w_c\| \cdot (\mathcal{F} \odot \mathcal{S}), where ⊙\odot is the Hadamard (element-wise) product.

    In contrast, the classification logit for class cc is computed using the global average pooled feature f(x)=GAP(F(x))∈RDf(x) = \text{GAP}(F(x)) \in \mathbb{R}^D:

    logitc(x)=wc⋅f(x)=∥wc∥∥f(x)∥S(wc,f(x))\text{logit}_c(x) = w_c \cdot f(x) = \|w_c\| \|f(x)\| S(w_c, f(x))

    Standard classification training optimizes alignment between wcw_c and the spatially averaged feature f(x)f(x), but does not enforce alignment between wcw_c and individual spatial feature vectors Fu(x)F_u(x). As a result, less discriminative object regions often have high feature norms Fu\mathcal{F}_u but low directional similarities Su\mathcal{S}_u, causing them to be suppressed in the final CAM.

  2. Knowl 2 — Feature Direction Alignment Objective via Similarity and Norm Losses

    model/method

    Feature direction alignment enforces directional alignment between spatial feature vectors Fu(x)∈RDF_u(x) \in \mathbb{R}^D and class-specific weight vector wc∈RDw_c \in \mathbb{R}^D in foreground regions while suppressing it in background regions via two complementary loss functions:

    1. Similarity Loss (Lsim\mathcal{L}_{sim}): Using the min-max normalized feature norm map F^=F−min⁡iFimax⁡iFi−min⁡iFi∈[0,1]H×W\hat{\mathcal{F}} = \frac{\mathcal{F} - \min_i \mathcal{F}_i}{\max_i \mathcal{F}_i - \min_i \mathcal{F}_i} \in [0, 1]^{H \times W}, coarse foreground and background regions are partitioned by constant thresholds τfg\tau_{fg} and τbg\tau_{bg}:

    Rfgnorm={u∣F^u>τfg},Rbgnorm={u∣F^u<τbg}R^{norm}_{fg} = \{u \mid \hat{\mathcal{F}}_u > \tau_{fg}\}, \quad R^{norm}_{bg} = \{u \mid \hat{\mathcal{F}}_u < \tau_{bg}\}

    The similarity loss maximizes cosine similarity Su=S(wc,Fu(x))\mathcal{S}_u = S(w_c, F_u(x)) in RfgnormR^{norm}_{fg} and minimizes it in RbgnormR^{norm}_{bg}:

    Lsim=−1∣Rfgnorm∣∑u∈RfgnormSu+1∣Rbgnorm∣∑u∈RbgnormSu\mathcal{L}_{sim} = -\frac{1}{|R^{norm}_{fg}|} \sum_{u \in R^{norm}_{fg}} \mathcal{S}_u + \frac{1}{|R^{norm}_{bg}|} \sum_{u \in R^{norm}_{bg}} \mathcal{S}_u

    1. Norm Loss (Lnorm\mathcal{L}_{norm}): To expand the normalized norm map F^\hat{\mathcal{F}} in candidate object areas that currently have low activation, candidate foreground and background regions are partitioned by the sign of the cosine similarity Su\mathcal{S}_u:

    Rfgsim={u∣Su>0},Rbgsim={u∣Su<0}R^{sim}_{fg} = \{u \mid \mathcal{S}_u > 0\}, \quad R^{sim}_{bg} = \{u \mid \mathcal{S}_u < 0\}

    (For fine-grained recognition where category objects share similar morphology across classes, RbgsimR^{sim}_{bg} is defined as regions having non-positive cosine similarity with all classes).

    The norm loss enforces high normalized feature norms in RfgsimR^{sim}_{fg} and low norms in RbgsimR^{sim}_{bg}:

    Lnorm=−1∣Rfgsim∣∑u∈RfgsimF^u+1∣Rbgsim∣∑u∈RbgsimF^u\mathcal{L}_{norm} = -\frac{1}{|R^{sim}_{fg}|} \sum_{u \in R^{sim}_{fg}} \hat{\mathcal{F}}_u + \frac{1}{|R^{sim}_{bg}|} \sum_{u \in R^{sim}_{bg}} \hat{\mathcal{F}}_u

    Joint optimization of Lsim\mathcal{L}_{sim} and Lnorm\mathcal{L}_{norm} drives the spatial activations of F^\hat{\mathcal{F}} and S\mathcal{S} to coincide over the full object extent.

  3. Knowl 3 — Consistency with Attentive Dropout

    model/method

    To ensure that normalized feature norms F^\hat{\mathcal{F}} are evenly distributed across both primary and secondary discriminative parts of the target object, consistency with attentive dropout applies a stochastic feature erasing mechanism paired with an L1L_1 representation consistency loss.

    Given an intermediate feature map F′∈RH′×W′×D′F' \in \mathbb{R}^{H' \times W' \times D'}, the channel-averaged activation map Mavg∈RH′×W′M^{avg} \in \mathbb{R}^{H' \times W'} is computed at each spatial location uu as:

    Muavg=1D′∑d=1D′Fu,d′M^{avg}_u = \frac{1}{D'} \sum_{d=1}^{D'} F'_{u, d}

    A threshold mask M1∈{0,1}H′×W′M_1 \in \{0, 1\}^{H' \times W'} identifies peak discriminative spatial locations:

    M1,u=I(Muavg>γ⋅max⁡kMkavg)M_{1, u} = \mathbb{I}\left(M^{avg}_u > \gamma \cdot \max_k M^{avg}_k\right)

    where γ∈(0,1)\gamma \in (0, 1) is an activation threshold parameter. A stochastic drop mask M2∈{0,1}H′×W′M_2 \in \{0, 1\}^{H' \times W'} samples independent Bernoulli random variables with drop probability pp:

    M2,u∼Bernoulli(p)M_{2, u} \sim \text{Bernoulli}(p)

    The masked intermediate feature map Fdrop′F'_{drop} is constructed as:

    Fdrop′=(1−M1⊙M2)⊙F′F'_{drop} = (1 - M_1 \odot M_2) \odot F'

    Let F(x)∈RH×W×DF(x) \in \mathbb{R}^{H \times W \times D} be the final feature map produced by feedforwarding F′F' through the remaining layers, and let Fdrop(x)∈RH×W×DF_{drop}(x) \in \mathbb{R}^{H \times W \times D} be the final feature map produced by feedforwarding Fdrop′F'_{drop} through the identical remaining layers with shared weights. The attentive dropout consistency loss is:

    Ldrop=∥F(x)−Fdrop(x)∥1\mathcal{L}_{drop} = \|F(x) - F_{drop}(x)\|_1

    This forces the network to preserve the complete object representation even when the most salient discriminative features are dropped.

  4. Knowl 4 — Two-Stage Training Scheme for Feature Alignment WSOL

    model/method

    The overall multi-task training objective combines standard classification cross-entropy loss LCE\mathcal{L}_{CE}, attentive dropout consistency loss Ldrop\mathcal{L}_{drop}, similarity loss Lsim\mathcal{L}_{sim}, and norm loss Lnorm\mathcal{L}_{norm}:

    Ltotal=LCE+λdropLdrop+λsimLsim+λnormLnorm\mathcal{L}_{total} = \mathcal{L}_{CE} + \lambda_{drop}\mathcal{L}_{drop} + \lambda_{sim}\mathcal{L}_{sim} + \lambda_{norm}\mathcal{L}_{norm}

    where λdrop\lambda_{drop}, λsim\lambda_{sim}, and λnorm\lambda_{norm} are hyperparameter balancing coefficients.

    Because feature direction alignment requires feature maps and class weights to develop initial discriminative structure, training is divided into two phases:

    1. Warm Stage (Initial Epochs): The model is trained exclusively with classification loss and attentive dropout consistency: Lwarm=LCE+λdropLdrop\mathcal{L}_{warm} = \mathcal{L}_{CE} + \lambda_{drop}\mathcal{L}_{drop}

    2. Joint Alignment Stage (Remaining Epochs): The model is optimized using the complete loss function Ltotal\mathcal{L}_{total}, introducing Lsim\mathcal{L}_{sim} and Lnorm\mathcal{L}_{norm} to align feature directions across the full object region.

    Standard default hyperparameter configurations are λdrop=3.0\lambda_{drop} = 3.0, λsim=0.5\lambda_{sim} = 0.5, λnorm=0.15\lambda_{norm} = 0.15, τfg=0.6\tau_{fg} = 0.6, τbg=0.1\tau_{bg} = 0.1, γ=0.8\gamma = 0.8, and p=0.5p = 0.5.

  5. Knowl 5 — WSOL Benchmark Performance on CUB-200-2011 and ImageNet-1K

    data/table

    Weakly supervised object localization accuracy evaluated on the CUB-200-2011 test set (5,794 images, 200 classes) and ImageNet-1K validation set (50,000 images, 1,000 classes) across VGG16 and ResNet50 backbones. Metrics reported are Top-1 Localization Accuracy (Top-1 Loc, %), Top-5 Localization Accuracy (Top-5 Loc, %), and Ground-Truth Localization Accuracy (GT Loc, %, evaluated with ground truth class at IoU≥0.5\text{IoU} \ge 0.5):

    Dataset Method / Backbone Top-1 Loc Top-5 Loc GT Loc
    CUB-200-2011 VGG16: CAM (CVPR '16) 44.15 52.16 56.00
    VGG16: ADL (CVPR '19) 52.36 - 75.41
    VGG16: DANet (ICCV '19) 52.52 61.96 67.70
    VGG16: DGL (ACMMM '20) 56.07 68.50 74.63
    VGG16: Bae et al. (ECCV '20) 58.96 - 76.30
    VGG16: Pan et al. (CVPR '21) 60.27 72.45 77.29
    VGG16: FAM (ICCV '21, multi-branch) 69.26 - 89.26
    VGG16: Ours (single-branch) 70.83 88.07 93.17
    ResNet50: CAM (CVPR '16) 46.91 53.57 -
    ResNet50: ADL (CVPR '19) 57.40 - 71.99
    ResNet50: DGL (ACMMM '20) 60.82 70.50 74.65
    ResNet50: Bae et al. (ECCV '20) 59.53 - 77.58
    ResNet50: Ours 73.16 86.68 91.60
    ImageNet-1K VGG16: CAM (CVPR '16) 42.80 54.86 -
    VGG16: ACoL (CVPR '18) 45.83 59.43 62.96
    VGG16: I2C\text{I}^2\text{C} (ECCV '20) 47.41 58.51 63.90
    VGG16: DGL (ACMMM '20) 47.66 58.89 64.78
    VGG16: Pan et al. (CVPR '21) 49.56 61.32 65.05
    VGG16: Ours 49.94 63.25 68.92
    ResNet50: ADL (CVPR '19) 48.23 - 61.04
    ResNet50: Bae et al. (ECCV '20) 49.42 - 62.20
    ResNet50: I2C\text{I}^2\text{C} (ECCV '20) 54.83 64.60 68.50
    ResNet50: DGL (ACMMM '20) 53.41 62.69 69.34
    ResNet50: Ours 53.76 65.75 69.89

    The feature direction alignment method achieves state-of-the-art localization across benchmarks without requiring auxiliary localization branches or multi-stage network cascades. On CUB-200-2011 with VGG16, it exceeds the prior best single-branch CAM method (Pan et al.) by 10.56%p in Top-1 Loc and 15.88%p in GT Loc, and surpasses the multi-branch method FAM by 1.57%p in Top-1 Loc and 3.91%p in GT Loc.

  6. Knowl 6 — Strict Intersection over Union Evaluation via MaxBoxAccV2

    data/table

    Localization performance under the MaxBoxAccV2 evaluation protocol on CUB-200-2011 and ImageNet-1K across IoU thresholds δ∈{0.3,0.5,0.7}\delta \in \{0.3, 0.5, 0.7\} and their arithmetic mean:

    CUB-200-2011 (VGG16) ImageNet-1K (VGG16)
    Method δ=0.3\delta=0.3 δ=0.5\delta=0.5 δ=0.7\delta=0.7 Mean δ=0.3\delta=0.3 δ=0.5\delta=0.5 δ=0.7\delta=0.7 Mean
    CAM 96.8 73.1 21.2 63.7 81.0 62.0 37.1 60.0
    HaS 92.1 69.9 29.1 63.7 80.7 62.1 38.9 60.6
    SPG 90.5 61.0 17.4 56.3 81.4 62.0 36.3 59.9
    ADL 97.7 78.1 23.0 66.3 80.8 60.9 37.8 59.9
    CutMix 91.1 67.3 28.6 62.3 80.3 61.0 37.1 59.5
    Ki et al. 96.2 77.2 26.8 66.7 81.5 63.2 39.4 61.3
    HaS + PaS - - - 61.2 - - - 62.1
    CALM - - - 64.8 - - - 62.8
    ADL + IVR - - - 71.5 - - - 63.7
    Ours 99.3 93.2 47.8 80.1 84.8 69.2 45.9 66.6
    CUB-200-2011 (ResNet50) ImageNet-1K (ResNet50)
    Method δ=0.3\delta=0.3 δ=0.5\delta=0.5 δ=0.7\delta=0.7 Mean δ=0.3\delta=0.3 δ=0.5\delta=0.5 δ=0.7\delta=0.7 Mean
    CAM 95.7 73.3 19.9 63.0 83.7 65.7 41.6 63.7
    HaS 93.1 72.2 28.6 64.6 83.7 65.2 41.3 63.4
    ADL 91.8 64.8 18.4 58.3 83.6 65.6 41.8 63.7
    Ki et al. 96.2 72.8 20.6 63.2 84.3 67.6 43.6 65.2
    CALM - - - 71.0 - - - 63.4
    ADL + IVR - - - 67.1 - - - 65.1
    Ours 99.4 90.4 38.0 75.9 86.7 71.1 48.3 68.7

    The feature direction alignment method delivers large improvements under strict IoU constraints (δ=0.7\delta = 0.7), gaining 21.0%p over Ki et al. with VGG16 (from 26.8% to 47.8%) and 17.4%p with ResNet50 (from 20.6% to 38.0%) on CUB-200-2011, demonstrating tight alignment with true object boundaries.

  7. Knowl 7 — Ablation Study of Feature Direction Alignment and Dropout Losses

    empirical result

    Ablation of the three contributed loss terms on the CUB-200-2011 test set using VGG16:

    Ldrop\mathcal{L}_{drop} Lsim\mathcal{L}_{sim} Lnorm\mathcal{L}_{norm} Top-1 Loc (%) Top-5 Loc (%) GT Loc (%)
    46.95 57.23 60.74
    ✓ 54.35 70.37 75.06
    ✓ 56.66 71.38 76.10
    ✓ ✓ 62.27 77.48 81.93
    ✓ ✓ 63.00 79.93 85.35
    ✓ ✓ ✓ 70.83 88.07 93.17

    Key observations:

    1. Attentive dropout consistency (Ldrop\mathcal{L}_{drop}) alone raises Top-1 Loc by 7.40%p and GT Loc by 14.32%p relative to baseline.
    2. Similarity loss (Lsim\mathcal{L}_{sim}) alone produces the largest single-component gain (+9.71%p Top-1 Loc, +15.36%p GT Loc).
    3. Adding the norm loss (Lnorm\mathcal{L}_{norm}) to Lsim\mathcal{L}_{sim} provides a +5.61%p gain in Top-1 Loc and +5.83%p gain in GT Loc.
    4. Combining all three loss terms achieves the peak performance of 70.83% Top-1 Loc and 93.17% GT Loc (+23.88%p Top-1 Loc and +32.43%p GT Loc over the baseline).
  8. Knowl 8 — Localization Capacity of Individual Decomposed Maps

    empirical result

    When trained with the full objective ({LCE,Ldrop,Lsim,Lnorm}\{L_{CE}, L_{drop}, L_{sim}, L_{norm}\}) on CUB-200-2011 using VGG16, bounding boxes predicted separately from the norm map F\mathcal{F} or the cosine similarity map S\mathcal{S} yield localization performance nearly identical to those extracted from the full CAM=∥wc∥⋅(F⊙S)\text{CAM} = \|w_c\| \cdot (\mathcal{F} \odot \mathcal{S}):

    Localization Map Top-1 Loc (%) Top-5 Loc (%) GT Loc (%)
    CAM\text{CAM} 70.83 88.07 93.17
    F\mathcal{F} (Feature Norm Map) 69.90 86.68 91.96
    S\mathcal{S} (Cosine Similarity Map) 70.38 87.64 93.13

    In standard CAM models, the similarity map S\mathcal{S} fails to localize objects because features in less discriminative parts are misaligned with wcw_c. Feature direction alignment brings F\mathcal{F} and S\mathcal{S} into spatial agreement, such that either decomposed component individually contains sufficient localization cues to identify the entire object.

  9. Knowl 9 — Performance Comparison of Attentive Dropout Consistency Against Erasing Integrated Learning

    empirical result

    On the CUB-200-2011 dataset with VGG16, combining feature direction alignment (Lsim+Lnorm\mathcal{L}_{sim} + \mathcal{L}_{norm}) with attentive dropout consistency (Ldrop\mathcal{L}_{drop}) outperforms combining feature direction alignment with Erasing Integrated Learning (EIL):

    Method Top-1 Loc (%) Top-5 Loc (%) GT Loc (%)
    Feature Direction Alignment only 62.27 77.48 81.93
    EIL + Feature Direction Alignment 66.10 82.21 86.78
    Attentive Dropout Consistency + Feature Direction Alignment 70.83 88.07 93.17

    While EIL trains the classifier to predict the correct label from partially erased features, attentive dropout consistency directly penalizes the L1L_1 difference ∥F(x)−Fdrop(x)∥1\|F(x) - F_{drop}(x)\|_1 between the full and masked intermediate representations. This explicit representation constraint distributes feature activation more uniformly across the object, yielding a 4.73%p higher Top-1 Loc and 6.39%p higher GT Loc over EIL + Alignment.

  10. Knowl 10 — Hyperparameter Sensitivity in Feature Alignment and Dropout Losses

    empirical result

    Empirical sensitivity evaluation of the hyperparameters on CUB-200-2011 with VGG16:

    1. Loss Weights (λsim,λnorm,λdrop\lambda_{sim}, \lambda_{norm}, \lambda_{drop}):

      • The objective is most sensitive to the similarity loss weight λsim\lambda_{sim}, with peak performance at λsim=0.5\lambda_{sim} = 0.5.
      • The norm loss weight λnorm\lambda_{norm} achieves optimal performance at λnorm=0.15\lambda_{norm} = 0.15 with minimal performance fluctuation across nearby values.
      • The dropout weight λdrop\lambda_{drop} achieves peak performance at λdrop=3.0\lambda_{drop} = 3.0; setting λdrop≥4.0\lambda_{drop} \ge 4.0 creates an overly rigid constraint that degrades GT Loc.
    2. Region Thresholds (τfg,τbg\tau_{fg}, \tau_{bg}):

      • Setting foreground threshold τfg=0.6\tau_{fg} = 0.6 and background threshold τbg=0.1\tau_{bg} = 0.1 for Lsim\mathcal{L}_{sim} reliably segments coarse object regions. GT Loc is stable across τfg∈[0.4,0.8]\tau_{fg} \in [0.4, 0.8] and τbg∈[0.06,0.14]\tau_{bg} \in [0.06, 0.14].
    3. Dropout Parameters (γ,p\gamma, p):

      • Activation threshold factor γ∈[0.7,0.9]\gamma \in [0.7, 0.9] is stable, with γ=0.8\gamma = 0.8 optimal; dropping too broadly at γ=0.6\gamma = 0.6 lowers localization performance.
      • Drop probability pp performs stably for stochastic dropout across p∈[0.25,0.75]p \in [0.25, 0.75] (optimal at p=0.5p = 0.5). In contrast, setting deterministic dropout p=1.0p = 1.0 causes a sharp collapse in localization accuracy, demonstrating that retaining partial discriminative cues during feature masking is necessary.

Coverage note — None was omitted; all key theoretical definitions, loss formulations, dropout and training procedures, benchmark tables (CUB-200-2011, ImageNet-1K, MaxBoxAccV2), ablation experiments, map coincidence results, and hyperparameter sensitivity analyses from the paper are fully covered.

References

  1. 1.Wonho Bae, Junhyug Noh, and Gunhee Kim. Rethinking class activation mapping for weakly supervised object localization. In European Conference on Computer Vision, pages 618–634. Springer, 2020.
  2. 2.Junsuk Choe, Seungho Lee, and Hyunjung Shim. Attention-based dropout layer for weakly supervised single object localization and semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020.
  3. 3.Junsuk Choe, Seong Joon Oh, Sanghyuk Chun, Zeynep Akata, and Hyunjung Shim. Evaluation for weakly supervised object localization: Protocol, metrics, and datasets. arXiv preprint arXiv:2007.04178, 2020.
  4. 4.Junsuk Choe and Hyunjung Shim. Attention-based dropout layer for weakly supervised object localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2219–2228, 2019.
  5. 5.Kaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi, Qingming Huang, and Qi Tian. Centernet: Keypoint triplets for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6569–6578, 2019.
  6. 6.Guangyu Guo, Junwei Han, Fang Wan, and Dingwen Zhang. Strengthen learning tolerance for weakly supervised object localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7403–7412, 2021.
  7. 7.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  8. 8.Minsong Ki, Youngjung Uh, Wonyoung Lee, and Hyeran Byun. In-sample contrastive learning and consistent attention for weakly supervised object localization. In Proceedings of the Asian Conference on Computer Vision, 2020.
  9. 9.Jeesoo Kim, Junsuk Choe, Sangdoo Yun, and Nojun Kwak. Normalization matters in weakly supervised object localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021.
  10. 10.Jae Myung Kim, Junsuk Choe, Zeynep Akata, and Seong Joon Oh. Keep calm and improve visual feature attribution. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021.
  11. 11.Jungbeom Lee, Eunji Kim, and Sungroh Yoon. Anti-adversarially manipulated attributions for weakly and semi-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4071–4080, 2021.
  12. 12.Min Lin, Qiang Chen, and Shuicheng Yan. Network in network. arXiv preprint arXiv:1312.4400, 2013.
  13. 13.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2117–2125, 2017.
  14. 14.Weizeng Lu, Xi Jia, Weicheng Xie, Linlin Shen, Yicong Zhou, and Jinming Duan. Geometry constrained weakly supervised object localization. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVI 16, pages 481–496. Springer, 2020.
  15. 15.Jinjie Mai, Meng Yang, and Wenfeng Luo. Erasing integrated learning: A simple yet effective approach for weakly supervised object localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8766–8775, 2020.
  16. 16.Meng Meng, Tianzhu Zhang, Qi Tian, Yongdong Zhang, and Feng Wu. Foreground activation maps for weakly supervised object localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3385–3395, 2021.
  17. 17.Xingjia Pan, Yingguo Gao, Zhiwen Lin, Fan Tang, Weiming Dong, Haolei Yuan, Feiyue Huang, and Changsheng Xu. Unveiling the potential of structure preserving for weakly supervised object localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11642–11651, 2021.
  18. 18.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28:91–99, 2015.
  19. 19.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015.
  20. 20.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  21. 21.Krishna Kumar Singh and Yong Jae Lee. Hide-and-seek: Forcing a network to be meticulous for weakly-supervised object and action localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3544–3553. IEEE, 2017.
  22. 22.Chuangchuang Tan, Guanghua Gu, Tao Ruan, Shikui Wei, and Yao Zhao. Dual-gradients localization framework for weakly supervised object localization. In Proceedings of the 28th ACM International Conference on Multimedia, pages 1976–1984, 2020.
  23. 23.Mingxing Tan, Ruoming Pang, and Quoc V Le. Efficientdet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10781–10790, 2020.
  24. 24.Jun Wei, Qin Wang, Zhen Li, Sheng Wang, S Kevin Zhou, and Shuguang Cui. Shallow feature matters for weakly supervised object localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5993–6001, 2021.
  25. 25.Peter Welinder, Steve Branson, Takeshi Mita, Catherine Wah, Florian Schroff, Serge Belongie, and Pietro Perona. Caltech-UCSD birds 200. Technical Report CNS-TR-2010-001, California Institute of Technology, 2010.
  26. 26.Jinheng Xie, Cheng Luo, Xiangping Zhu, Ziqi Jin, Weizeng Lu, and Linlin Shen. Online refinement of low-level feature based activation map for weakly supervised object localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 132–141, 2021.
  27. 27.Haolan Xue, Chang Liu, Fang Wan, Jianbin Jiao, Xiangyang Ji, and Qixiang Ye. Danet: Divergent activation for weakly supervised object localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6589–6598, 2019.
  28. 28.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6023–6032, 2019.
  29. 29.Chen-Lin Zhang, Yun-Hao Cao, and Jianxin Wu. Rethinking the route towards weakly supervised object localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13460–13469, 2020.
  30. 30.Xiaolin Zhang, Yunchao Wei, Jiashi Feng, Yi Yang, and Thomas S Huang. Adversarial complementary learning for weakly supervised object localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1325–1334, 2018.
  31. 31.Xiaolin Zhang, Yunchao Wei, Guoliang Kang, Yi Yang, and Thomas Huang. Self-produced guidance for weakly-supervised object localization. In Proceedings of the European conference on computer vision (ECCV), pages 597–613, 2018.
  32. 32.Xiaolin Zhang, Yunchao Wei, and Yi Yang. Inter-image communication for weakly supervised localization. In European Conference on Computer Vision, pages 271–287. Springer, 2020.
  33. 33.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2921–2929, 2016.

Citation

MLA
Kim, E., et al. “Bridging the Gap Between Classification and Localization for Weakly Supervised Object Localization”. arXiv, 2022, http://arxiv.org/abs/2204.00220v1.
APA
Kim, E., Kim, S., Lee, J., Kim, H., & Yoon, S. (2022). Bridging the Gap between Classification and Localization for Weakly Supervised Object Localization. arXiv. http://arxiv.org/abs/2204.00220v1
Chicago
Kim, E., S. Kim, J. Lee, H. Kim, and S. Yoon. 2022. “Bridging the Gap Between Classification and Localization for Weakly Supervised Object Localization”. arXiv. http://arxiv.org/abs/2204.00220v1.
Harvard
Kim, E. et al. (2022) “Bridging the Gap between Classification and Localization for Weakly Supervised Object Localization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.00220v1.
Vancouver
1. Kim E, Kim S, Lee J, Kim H, Yoon S (2022) Bridging the Gap between Classification and Localization for Weakly Supervised Object Localization. arXiv

BibTeX

@article{kim2022bridging,
  title = {Bridging the Gap between Classification and Localization for Weakly Supervised Object Localization},
  author = {Kim, Eunji and Kim, Siwon and Lee, Jungbeom and Kim, Hyunwoo and Yoon, Sungroh},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.00220v1},
  eprint = {2204.00220}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE