SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning

Hongjun WangSagar VazeKai Han

article2024ICLR73 citations

Introduces SPTNet, an efficient Generalized Category Discovery framework that alternates model fine-tuning with spatial prompt tuning to achieve a 10% accuracy gain over prior state-of-the-art baselines on the Semantic Shift Benchmark with only 0.117% parameter overhead.

Listen

Real-world computer vision systems frequently encounter open-world scenarios where incoming data includes both previously seen categories and completely novel, unseen categories. Generalized category discovery addresses this challenge by using partially labelled data from known categories to classify unlabelled images across both seen and novel classes. Standard approaches typically fine-tune large pre-trained vision models, which introduces significant computational costs and risks overfitting to the labelled data.

The main objective of the article is to demonstrate SPTNet, a highly parameter-efficient learning framework that integrates spatial prompt tuning to categorize unlabelled images across known and unknown classes without heavy model retraining.

The evaluated approach divides input images into local patches and injects learnable, pixel-level prompt parameters around each patch as well as a global border around the entire image. Instead of jointly optimizing the model and prompts—which leads to optimization instability—the framework alternates in two stages: freezing the vision backbone to update prompt parameters, and then freezing the prompts to fine-tune only the projection head and the top layer of the vision backbone using contrastive learning. This alternating optimization treats spatial prompts as targeted, learned data augmentations. The article evaluates this method across seven standard image benchmarks, spanning generic datasets (such as CIFAR-10, CIFAR-100, and ImageNet-100) and fine-grained benchmarks (such as CUB, Stanford Cars, FGVC-Aircraft, and Herbarium-19).

The article demonstrates several key findings:

  1. On the challenging fine-grained Semantic Shift Benchmark, the method achieves an average accuracy of 61.4%, surpassing prior state-of-the-art baselines by approximately 10% in proportional terms (and about 5% in absolute terms).
  2. Across generic image datasets, the approach consistently matches or exceeds existing methods, reaching 97.3% accuracy on CIFAR-10, 81.3% on CIFAR-100, and 85.4% on ImageNet-100.
  3. The full framework achieves these gains while introducing only 0.117% additional parameters relative to the base Vision Transformer architecture, with simpler patch-level variants requiring as little as 0.039% extra parameters.
  4. The alternating two-stage training scheme substantially outperforms standard end-to-end joint training, which causes prompts to degrade and become inactive, while cutting training time roughly in half compared to competing discovery methods.

These results show that adapting the input data representation through local spatial prompting is both more effective and far more resource-efficient than extensively retraining large backbone models. By directing the visual model's attention toward critical, fine-grained object parts, the framework enables strong knowledge transfer from known to unknown categories. This substantially reduces computing requirements and training turnaround times for deploying vision systems in dynamic environments.

Decision-makers and practitioners implementing open-world vision pipelines should adopt parameter-efficient spatial prompting and alternating optimization over full model fine-tuning. For operational deployment, smaller prompt sizes (such as a single-pixel border per patch) should be selected to avoid occluding core visual content. Future work should focus on validating the approach in production pilot environments and exploring more advanced foundation models like DINOv2 to further improve novel category discovery.

The framework exhibits certain limitations. Performance gains are less pronounced on low-resolution images (such as 32x32 pixels), where image patches lack sufficient spatial detail. Additionally, performance declines when models face significant cross-domain shifts or when the true category count is unknown, and the underlying decision-making of the learned prompts lacks full interpretability. Nonetheless, the consistent empirical improvements across multiple benchmarks provide high confidence in the framework's effectiveness for fine-grained open-world visual discovery.

  • Paper: Generalized Category Discovery with Decoupled Prototypical Network, Wenbin An et al. (2023). This paper establishes the generalized category discovery formulation and prototype-based transfer methods that SPTNet directly aims to improve via efficient spatial prompt tuning.
  • Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). This work introduces visual prompt tuning as a parameter-efficient alternative to full fine-tuning of vision Transformers, providing the conceptual foundation for SPTNet's spatial prompt design.
  • Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). This study introduces context optimization and prompt tuning paradigms for adapting frozen foundation models, establishing the parameter-efficient adaptation principles leveraged by SPTNet.
  • Paper: DINOv2: Learning Robust Visual Features without Supervision, Maxime Oquab et al. (2023). This paper provides key insights into self-supervised Vision Transformers and patch-level feature learning that underpin modern backbones and baseline representations evaluated in generalized category discovery.

No sufficiently relevant recommendations were found.

Cover for SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning

Abstract

Generalized Category Discovery (GCD) aims to classify unlabelled images from both seen' and unseen' classes by transferring knowledge from a set of labelled `seen' class images. A key theme in existing GCD approaches is adapting large-scale pre-trained models for the GCD task. An alternate perspective, however, is to adapt the data representation itself for better alignment with the pre-trained model. As such, in this paper, we introduce a two-stage adaptation approach termed SPTNet, which iteratively optimizes model parameters (i.e., model-finetuning) and data parameters (i.e., prompt learning). Furthermore, we propose a novel spatial prompt tuning method (SPT) which considers the spatial property of image data, enabling the method to better focus on object parts, which can transfer between seen and unseen classes. We thoroughly evaluate our SPTNet on standard benchmarks and demonstrate that our method outperforms existing GCD methods. Notably, we find our method achieves an average accuracy of 61.4% on the SSB, surpassing prior state-of-the-art methods by approximately 10%. The improvement is particularly remarkable as our method yields extra parameters amounting to only 0.117% of those in the backbone architecture. Project page: this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Methods
  • 3.1 Preliminaries
  • 3.2 SPTNet: an alternate prompt learning framework for GCD
  • 3.3 Spatial Prompt Tuning
  • 4 Experiments
  • 4.1 Experimental setup
  • 4.2 Main results
  • 4.3 Ablation study
  • 4.4 Qualitative comparison
  • 5 Conclusion
  • References
  • A Discussion on different variants of SPTNet
  • B Training configurations for SPTNet-P and SPTNet-S
  • C Benchmarking results of SPTNet-P and SPTNet-S
  • D Effects of global prompt size m+m^{+}
  • E Visualization of learned prompts
  • F Results based on DINOv2
  • G Robustness of SPTNet for GCD with domain shifts
  • H Unknown category number
  • I Performance and time efficiency
  • J Theoretical analysis on our alternate training
  • K More visualization and analysis of attention maps
  • L Broader impacts

Knowls

  1. Knowl 1 — Two-Stage Alternating Framework (SPTNet) for Generalized Category Discovery

    model/method

    In Generalized Category Discovery (GCD), a model is trained on a dataset D\mathcal{D} consisting of a labeled subset Dl={(Xi,yi)}i=1Nl⊂Xl×Yl\mathcal{D}_l = \{(X_i, y_i)\}_{i=1}^{N_l} \subset \mathcal{X}_l \times \mathcal{Y}_l from seen classes C1\mathcal{C}_1 and an unlabeled subset Du={Xi}i=1Nu⊂Xu\mathcal{D}_u = \{X_i\}_{i=1}^{N_u} \subset \mathcal{X}_u drawn from both seen and unseen classes C=C1∪C2\mathcal{C} = \mathcal{C}_1 \cup \mathcal{C}_2. The target is to categorize all images in Du\mathcal{D}_u.

    SPTNet models the prediction using a Vision Transformer (ViT) feature extractor F\mathcal{F} followed by a multilayer perceptron (MLP) projection head H\mathcal{H}, parameterized as y^=H(F(X))\hat{y} = \mathcal{H}(\mathcal{F}(X)). Instead of jointly optimizing model parameters and prompt tokens (which causes training instability and trivial shortcut representations), SPTNet decouples optimization into two alternating stages across kk-iteration intervals:

    1. Stage 1 (Data Parameter Tuning): The parameters of the backbone F\mathcal{F} and projection head H\mathcal{H} are frozen. Only spatial prompt parameters Ps\mathcal{P}_s attached to patch inputs are updated using the combined contrastive and classification loss. Weight decay for prompt optimization is set to zero to prevent sparsity and preserve perturbation diversity.
    2. Stage 2 (Model Parameter Tuning): Prompt parameters Ps\mathcal{P}_s are fixed. The prompted inputs ϕ(X)+Ps\phi(X) + \mathcal{P}_s serve as a dynamic, learned data augmentation. The projection head H\mathcal{H} and the final Transformer block of F\mathcal{F} are updated using the objective function, forcing the backbone to learn invariant representations against the learned spatial perturbations.
  2. Knowl 2 — Spatial Prompt Tuning (SPT) Formulation and Architectural Variants

    model/method

    Spatial Prompt Tuning (SPT) adapts pre-trained Vision Transformers (ViTs) by attaching pixel-level learnable prompts directly to local image patches rather than injecting continuous tokens into hidden transformer layers or wrapping only the outer image border.

    Given an image X∈R3×H×WX \in \mathbb{R}^{3 \times H \times W} partitioned by a patchify operator ϕ(X)=(x1,x2,…,xn)\phi(X) = (x^1, x^2, \dots, x^n) into n=(H×W)/(h×w)n = (H \times W)/(h \times w) patches where xj∈R3×h×wx^j \in \mathbb{R}^{3 \times h \times w}, SPT learns an instance-agnostic spatial prompt tensor Ps={p1,p2,…,pn}\mathcal{P}_s = \{p^1, p^2, \dots, p^n\}. For each patch xjx^j, prompt values pj(c,d)∈R3p^j(c, d) \in \mathbb{R}^3 at pixel coordinates c∈{0,…,h−1}c \in \{0, \dots, h-1\} and d∈{0,…,w−1}d \in \{0, \dots, w-1\} are constrained by a border width mm:

    pj(c,d)={0if m≤c<h−m and m≤d<w−mlearnable parameterotherwisep^j(c, d) = \begin{cases} 0 & \text{if } m \le c < h - m \text{ and } m \le d < w - m \\ \text{learnable parameter} & \text{otherwise} \end{cases}

    Each patch prompt contains 6m(h+w−2m)6m(h + w - 2m) parameters.

    Three prompt structural variants are defined:

    • SPTNet (Default): Attaches distinct spatial prompts to each patch and wraps an additional global prompt Pg\mathcal{P}_g of border width m+m^+ around the full image border. For H=W=224H=W=224, h=w=16h=w=16, m=1m=1, and m+=30m^+=30, it adds 105,120 extra parameters (0.117%0.117\% of ViT-Base).
    • SPTNet-P (Patch): Attaches distinct spatial prompts to each patch without a global prompt, totaling 6m(h+w−2m)n=35,2806m(h + w - 2m)n = 35,280 parameters (0.039%0.039\% of ViT-Base).
    • SPTNet-S (Shared): Shares a single spatial prompt across all nn patches without a global prompt, totaling 6m(h+w−2m)=1806m(h + w - 2m) = 180 parameters (0.0002%0.0002\% of ViT-Base).
  3. Knowl 3 — SPTNet Training Objective and Loss Formulation

    equation

    The overall objective function L\mathcal{L} for both Stage 1 and Stage 2 optimization in SPTNet combines unsupervised contrastive loss, supervised contrastive loss, unsupervised self-distillation classification loss, supervised classification loss, and a mean-entropy-maximization regularizer:

    L=(1−λ)(Lnceun+Lclsun)+λ(Lncesup+Lclssup)−ϵΔ\mathcal{L} = (1 - \lambda)(\mathcal{L}_{\text{nce}}^{\text{un}} + \mathcal{L}_{\text{cls}}^{\text{un}}) + \lambda(\mathcal{L}_{\text{nce}}^{\text{sup}} + \mathcal{L}_{\text{cls}}^{\text{sup}}) - \epsilon \Delta

    where λ∈[0,1]\lambda \in [0, 1] balances supervised and unsupervised objectives, and ϵ>0\epsilon > 0 balances the regularizer Δ\Delta.

    The components are defined as:

    1. Unsupervised InfoNCE Loss: Lnceun(X,X′,N(X);F,τu)=−log⁡exp⁡(cos⁡(F(X),F(X′))/τu)∑q=1Qexp⁡(cos⁡(F(X),F(Xq−))/τu)\mathcal{L}_{\text{nce}}^{\text{un}}(X, X', \mathcal{N}(X); \mathcal{F}, \tau_u) = -\log \frac{\exp(\cos(\mathcal{F}(X), \mathcal{F}(X')) / \tau_u)}{\sum_{q=1}^Q \exp(\cos(\mathcal{F}(X), \mathcal{F}(X_q^-)) / \tau_u)} where X,X′X, X' are two augmented views of an image, N(X)={Xq−}q=1Q\mathcal{N}(X) = \{X_q^-\}{q=1}^Q is a set of negative samples, cos⁡(⋅,⋅)\cos(\cdot, \cdot) is cosine similarity, and τu\tau_u is a temperature hyperparameter.

    2. Supervised Contrastive Loss: Lncesup(X,P(X),N(X),y;F,τc)\mathcal{L}_{\text{nce}}^{\text{sup}}(X, \mathcal{P}(X), \mathcal{N}(X), y; \mathcal{F}, \tau_c) generalizes InfoNCE by averaging over all positive pairs P(X)\mathcal{P}(X) sharing the ground-truth label yy in the mini-batch.

    3. Supervised Classification Loss: Lclssup(X,y;H,F,τs)=−∑κ=1∣C∣yκlog⁡exp⁡(cos⁡(H(F(X)),Wκ)/τs)∑κ′=1∣C∣exp⁡(cos⁡(H(F(X)),Wκ′)/τs)\mathcal{L}_{\text{cls}}^{\text{sup}}(X, y; \mathcal{H}, \mathcal{F}, \tau_s) = -\sum_{\kappa=1}^{|\mathcal{C}|} y_\kappa \log \frac{\exp(\cos(\mathcal{H}(\mathcal{F}(X)), W_\kappa) / \tau_s)}{\sum_{\kappa'=1}^{|\mathcal{C}|} \exp(\cos(\mathcal{H}(\mathcal{F}(X)), W_{\kappa'}) / \tau_s)} where H(F(X))\mathcal{H}(\mathcal{F}(X)) is the ℓ2\ell_2-normalized feature, WκW_\kappa is the learnable prototype vector for class κ\kappa, and τs\tau_s is the temperature.

    4. Unsupervised Classification Loss: Lclsun(X,X′;H,F,τt)\mathcal{L}_{\text{cls}}^{\text{un}}(X, X'; \mathcal{H}, \mathcal{F}, \tau_t) calculates cross-entropy between the prediction on XX and the soft pseudo-label generated by y^=H(F(X′))\hat{y} = \mathcal{H}(\mathcal{F}(X')) via self-distillation.

    5. Mean Entropy Maximization Regularizer: Δ\Delta is computed as the entropy of the mean prediction distribution across all samples in the mini-batch to prevent cluster collapse.

  4. Knowl 4 — SPTNet Alternating Optimization Algorithm

    algorithm

    The two-stage alternating optimization procedure updates data parameters (prompts) and model parameters (last ViT block and projection head) in alternating intervals of kk iterations:

    Input: Labeled set Dl\mathcal{D}_l, unlabeled set Du\mathcal{D}_u, ViT backbone F\mathcal{F} pre-trained with DINO, projection head H\mathcal{H}, prompt parameters Ps\mathcal{P}_s, alternating frequency kk, total training epochs EE, prompt learning rate lrplr_p, backbone learning rate lrblr_b, batch size BB, balancing factors λ,ϵ\lambda, \epsilon, temperatures τu,τc,τs,τt\tau_u, \tau_c, \tau_s, \tau_t.
    Output: Trained feature extractor F\mathcal{F}, projection head H\mathcal{H}, and spatial prompts Ps\mathcal{P}_s.
    Initialize prompt parameters Ps\mathcal{P}_s randomly and prototype vectors WW in H\mathcal{H}.
    for epoch = 1 to EE do
        for step = 1 to number_of_batches do
            Sample mini-batch B=Bl∪Bu\mathcal{B} = \mathcal{B}_l \cup \mathcal{B}_u from Dl\mathcal{D}_l and Du\mathcal{D}_u.
            Generate augmented views XX and X′X'.
            Construct prompted patch representations Xˉ=ϕ(X)+Ps\bar{X} = \phi(X) + \mathcal{P}_s and Xˉ′=ϕ(X′)+Ps\bar{X}' = \phi(X') + \mathcal{P}_s.
            Compute objective loss L\mathcal{L} on (Xˉ,Xˉ′)(\bar{X}, \bar{X}') using Eq. (4).
            if ⌊(step−1)/k⌋(mod2)==0\lfloor (\text{step} - 1) / k \rfloor \pmod 2 == 0 then
                // Stage 1: Update Prompts
                Freeze parameters of F\mathcal{F} and H\mathcal{H}.
                Compute gradients ∇PsL\nabla_{\mathcal{P}_s} \mathcal{L}.
                Update Ps←Ps−lrp∇PsL\mathcal{P}_s \leftarrow \mathcal{P}_s - lr_p \nabla_{\mathcal{P}_s} \mathcal{L} (with weight decay wdp=0wd_p = 0).
            else
                // Stage 2: Update Backbone Top Layer and Head
                Freeze prompt parameters Ps\mathcal{P}_s.
                Compute gradients ∇H,FlastL\nabla_{\mathcal{H}, \mathcal{F}_{\text{last}}} \mathcal{L}.
                Update H,Flast←(H,Flast)−lrb∇H,FlastL\mathcal{H}, \mathcal{F}_{\text{last}} \leftarrow (\mathcal{H}, \mathcal{F}_{\text{last}}) - lr_b \nabla_{\mathcal{H}, \mathcal{F}_{\text{last}}} \mathcal{L} (with weight decay wdbwd_b).
            end if
        end for
    end for
    return F,H,Ps\mathcal{F}, \mathcal{H}, \mathcal{P}_s
  5. Knowl 5 — Expectation-Maximization Theoretical Analysis of Alternating Training

    theoretical result

    The alternating optimization of model parameters θ\theta (in F\mathcal{F} and H\mathcal{H}) and data prompt parameters p1:np^{1:n} corresponds to an Expectation-Maximization (EM) lower-bound maximization of the data log-likelihood L(θ)=ln⁡P(X∣θ)\mathcal{L}(\theta) = \ln P(X \mid \theta).

    For a parameter state θt\theta_t, the log-likelihood difference is lower-bounded by Jensen's inequality as:

    L(θ)−L(θt)≥∑p1:nP(p1:n∣X,θt)ln⁡P(X∣p1:n,θ)P(p1:n∣θ)P(p1:n∣X,θt)P(X∣θt)≜H(θ∣θt)\mathcal{L}(\theta) - \mathcal{L}(\theta_t) \ge \sum_{p^{1:n}} P(p^{1:n} \mid X, \theta_t) \ln \frac{P(X \mid p^{1:n}, \theta) P(p^{1:n} \mid \theta)}{P(p^{1:n} \mid X, \theta_t) P(X \mid \theta_t)} \triangleq \mathcal{H}(\theta \mid \theta_t)

    Defining l(θ∣θt)=L(θt)+H(θ∣θt)l(\theta \mid \theta_t) = \mathcal{L}(\theta_t) + \mathcal{H}(\theta \mid \theta_t), the parameter update θt+1=arg⁡max⁡θl(θ∣θt)\theta_{t+1} = \arg\max_\theta l(\theta \mid \theta_t) reduces to:

    θt+1=arg⁡max⁡θEp1:n∣X,θt[ln⁡P(X,p1:n∣θ)]\theta_{t+1} = \arg\max_\theta \mathbb{E}_{p^{1:n} \mid X, \theta_t} \left[ \ln P(X, p^{1:n} \mid \theta) \right]

    Because prompt parameters p1:np^{1:n} depend jointly on XX and θt\theta_t, the factorized likelihood P(X∣p1:n,θt)P(p1:n∣θt)≠P(X,p1:n∣θt)P(X \mid p^{1:n}, \theta_t) P(p^{1:n} \mid \theta_t) \ne P(X, p^{1:n} \mid \theta_t). Consequently, standard end-to-end simultaneous gradient descent over θ\theta and p1:np^{1:n} does not optimize this EM objective and allows the model to learn shortcut invariances (causing prompt parameters to deactivate toward zero). Alternating optimization enforces active prompt perturbations that act as valid, non-collapsing data augmentations.

  6. Knowl 6 — Performance of SPTNet on Semantic Shift Benchmark (SSB) and Fine-Grained Datasets

    data/table

    SPTNet was evaluated on fine-grained category discovery using CUB, Stanford Cars, FGVC-Aircraft, and Herbarium-19. In all benchmarks, labeled sets Dl\mathcal{D}_l were formed by sampling 50%50\% of images from the seen classes C1\mathcal{C}_1, while the remaining seen and all novel class C2\mathcal{C}_2 images constituted the unlabeled set Du\mathcal{D}_u. Accuracy was measured via Hungarian optimal matching clustering accuracy (ACC) on All, Old (seen), and New (unseen) categories.

    CUB Stanford Cars FGVC-Aircraft Herbarium19
    Method All Old New All Old New All Old New All Old New
    k-means 34.3 38.9 32.1 12.8 10.6 13.8 12.9 12.9 12.8 13.0 12.2 13.4
    RankStats+ 33.3 51.6 24.2 28.3 61.8 12.1 27.9 55.8 12.8 27.9 55.8 12.8
    UNO+ 35.1 49.0 28.1 35.5 70.5 18.6 28.3 53.7 14.7 28.3 53.7 14.7
    GCD 51.3 56.6 48.7 39.0 57.6 29.9 45.0 41.1 46.9 35.4 51.0 27.0
    ORCA 36.3 43.8 32.6 31.9 42.2 26.9 31.6 32.0 31.4 20.9 30.9 15.5
    SimGCD 60.3 65.6 57.7 53.8 71.9 45.0 54.2 59.1 51.8 43.0 58.0 35.1
    DCCL 63.5 60.8 64.9 43.1 55.7 36.2 - - - - - -
    PromptCAL 62.9 64.4 62.1 50.2 70.1 40.6 52.2 52.2 52.3 37.0 52.0 28.9
    SPTNet (Ours) 65.8 68.8 65.1 59.0 79.2 49.3 59.3 61.8 58.1 43.4 58.7 35.2

    SPTNet achieves an average clustering accuracy of 61.4%61.4\% on the three SSB datasets (CUB, Stanford Cars, FGVC-Aircraft), yielding an absolute gain of ∼5%\sim 5\% and a proportional improvement of ∼10%\sim 10\% over prior state-of-the-art methods on 'All' categories.

  7. Knowl 7 — Performance of SPTNet on Generic Image Recognition Datasets

    data/table

    Evaluation of SPTNet on generic category discovery benchmarks using CIFAR-10, CIFAR-100, and ImageNet-100. Labeled sets contain 50%50\% of seen class data for CIFAR-10 and ImageNet-100, and 80%80\% for CIFAR-100.

    CIFAR-10 CIFAR-100 ImageNet-100
    Method All Old New All Old New All Old New
    k-means 83.6 85.7 82.5 52.0 52.2 50.8 72.7 75.5 71.3
    RankStats+ 46.8 19.2 60.5 58.2 77.6 19.3 37.1 61.6 24.8
    UNO+ 68.6 98.3 53.8 69.5 80.6 47.2 70.3 95.0 57.9
    GCD 91.5 97.9 88.2 73.0 76.2 66.5 74.1 89.8 66.3
    ORCA 96.9 95.1 97.8 74.2 82.1 67.2 79.2 93.2 72.1
    SimGCD 97.1 95.1 98.1 80.1 81.2 77.8 83.0 93.1 77.9
    DCCL 96.3 96.5 96.9 75.3 76.8 70.2 80.5 90.5 76.2
    PromptCAL 97.9 96.6 98.5 81.2 84.2 75.3 83.1 92.7 78.3
    SPTNet (Ours) 97.3 95.0 98.6 81.3 84.3 75.6 85.4 93.2 81.4

    SPTNet outperforms SimGCD by 0.4%0.4\% on CIFAR-10, 1.9%1.9\% on CIFAR-100, and 2.5%2.5\% on ImageNet-100 for 'All' classes. On CIFAR datasets (32×3232 \times 32 resolution), patch information is constrained by low pixel resolution, resulting in smaller relative gains than on high-resolution fine-grained datasets.

  8. Knowl 8 — Ablation Analysis of Visual Prompting Strategies and Optimization Schedules

    empirical result

    Ablation studies on the Semantic Shift Benchmark (SSB averaged across CUB, Stanford Cars, and FGVC-Aircraft) and ImageNet-100 demonstrate the relative impact of prompt designs and training schedules:

    1. Prompt Architecture Comparison on SSB (All / Old / New ACC %):

      • SimGCD Baseline: 56.1/65.5/51.556.1 / 65.5 / 51.5
      • SimGCD + VPT (Visual Prompt Tuning in hidden layers): 54.4/64.7/49.154.4 / 64.7 / 49.1 (performance drops due to representation misalignment in contrastive feature space)
      • SimGCD + Global Prompt (input border): 56.7/64.6/53.556.7 / 64.6 / 53.5
      • SimGCD + SPT (local patch prompts, single stage): 57.9/67.2/53.357.9 / 67.2 / 53.3
      • SPTNet with Shared Prompts + Alternating Training (SPTNet-S): 60.5/68.6/56.560.5 / 68.6 / 56.5
      • SPTNet with Patch Prompts + Alternating Training (SPTNet-P): 59.1/68.5/54.559.1 / 68.5 / 54.5
      • SPTNet with Shared & Global Prompts + Alternating Training: 60.9/69.0/57.360.9 / 69.0 / 57.3
      • Full SPTNet (Patch + Global Prompts + Alternating Training): 61.4/69.9/57.561.4 / 69.9 / 57.5
    2. Training Schedules Comparison on ImageNet-100 and SSB (All ACC %):

      • SimGCD baseline: 83.0%83.0\% (ImageNet-100) / 56.1%56.1\% (SSB)
      • SimGCD further fine-tuned: 84.3%84.3\% (ImageNet-100) / 57.0%57.0\% (SSB)
      • SPTNet end-to-end joint training: 84.1%84.1\% (ImageNet-100) / 58.6%58.6\% (SSB)
      • SPTNet data-parameters-first: 83.5%83.5\% (ImageNet-100) / 58.0%58.0\% (SSB)
      • SPTNet model-parameters-first: 84.8%84.8\% (ImageNet-100) / 59.2%59.2\% (SSB)
      • SPTNet alternating training (k=20k=20): 85.4%85.4\% (ImageNet-100) / 61.4%61.4\% (SSB)
    3. Prompt Hyperparameters: Performance peaks at spatial prompt width m=1m=1 (larger mm causes excessive object occlusion, dropping SSB accuracy to 22.5%22.5\% at m=8m=8) and global prompt width m+=30m^+=30.

  9. Knowl 9 — Robustness of SPTNet Under Domain Shifts and Unknown Category Counts

    empirical result

    SPTNet exhibits robustness in non-ideal GCD operational environments:

    1. GCD with Domain Shift on DomainNet (Clustering ACC %): Trained on labeled 'real' and unlabeled 'painting' images, and tested across seen and unseen domains:

      • Real domain (labeled source): SPTNet achieves 63.1%63.1\% All / 75.9%75.9\% Old / 56.4%56.4\% New, compared to SimGCD's 61.3/77.8/52.961.3 / 77.8 / 52.9.
      • Painting domain (unlabeled target): SPTNet achieves 39.2%39.2\% All / 43.1%43.1\% Old / 35.2%35.2\% New, compared to SimGCD's 34.5/35.6/33.534.5 / 35.6 / 33.5.
      • Unseen domains (quickdraw, sketch, infograph, clipart): SPTNet achieves 17.4%17.4\% All / 22.2%22.2\% Old / 13.6%13.6\% New, compared to SimGCD's 16.7/22.5/12.216.7 / 22.5 / 12.2.
    2. Unknown Category Count Estimation: When true class count ∣C∣|\mathcal{C}| is unavailable and estimated via semi-supervised k-means clustering:

      • CUB (Estimated ∣extAll∣/∣Old∣=231/109| ext{All}|/|\text{Old}| = 231/109 vs True 200/100200/100): SPTNet-P achieves 65.2%65.2\% All / 71.0%71.0\% Old / 62.3%62.3\% New, outperforming SimGCD (61.0/66.0/58.661.0 / 66.0 / 58.6).
      • ImageNet-100 (Estimated 231/109231/109 vs True 200/100200/100): SPTNet-P achieves 83.4%83.4\% All / 91.8%91.8\% Old / 74.6%74.6\% New, outperforming SimGCD (81.1/90.9/76.181.1 / 90.9 / 76.1).
      • When varying the class multiplier factor C′∈{0.1,0.5,1.0,2.0,10.0}C' \in \{0.1, 0.5, 1.0, 2.0, 10.0\}, accuracy degrades smoothly rather than abruptly.
  10. Knowl 10 — Limitations of SPTNet

    limitation

    The authors identify three primary limitations of the SPTNet framework:

    1. Decision Interpretability: The precise mechanisms by which spatial prompt perturbations steer open-world representation clustering and partitioning between novel and seen classes lack complete formal transparency.
    2. Cross-Domain Generalization Gap: Although SPTNet outperforms existing GCD baselines under domain shifts on DomainNet, absolute discovery accuracy on unseen domains remains low (17.4%17.4\% on out-of-distribution domains).
    3. Backbone Dependency and Bias Inheritance: The method relies on pre-trained self-supervised vision backbones (e.g., DINO/DINOv2) as initial representations, inheriting representational biases, domain sensitivity, and potential privacy issues embedded in foundation model pre-training.

Coverage note — No substantial contributed material was omitted; qualitative attention heatmaps and DINOv2 ablation tables are summarized within the existing empirical results and ablation knowls.

References

  1. 1.David Arthur and Sergei Vassilvitskii. k-means++: The advantages of careful seeding. Technical report, Stanford, 2006.
  2. 2.Mahmoud Assran, Mathilde Caron, Ishan Misra, Piotr Bojanowski, Florian Bordes, Pascal Vincent, Armand Joulin, Mike Rabbat, and Nicolas Ballas. Masked siamese networks for label-efficient learning. In ECCV, 2022.
  3. 3.Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan, and Phillip Isola. Exploring visual prompts for adapting large-scale models. arXiv preprint arXiv: 2203.17274, 2022.
  4. 4.David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A Raffel. Mixmatch: A holistic approach to semi-supervised learning. In NeurIPS, 2019.
  5. 5.Geunyeong Byeon and Pascal Van Hentenryck. Benders subproblem decomposition for bilevel problems with convex follower. INFORMS Journal on Computing, 2022.
  6. 6.Kaidi Cao, Maria Brbic, and Jure Leskovec. Open-world semi-supervised learning. In ICLR, 2022.
  7. 7.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, 2021.
  8. 8.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020a.
  9. 9.Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E Hinton. Big self-supervised models are strong semi-supervised learners. In NeurIPS, 2020b.
  10. 10.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020c.
  11. 11.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In ICCV, 2021.
  12. 12.Arthur P Dempster, Nan M Laird, and Donald B Rubin. Maximum likelihood from incomplete data via the em algorithm. Journal of the royal statistical society: series B (methodological), 1977.
  13. 13.Bowen Dong, Pan Zhou, Shuicheng Yan, and Wangmeng Zuo. Lpt: Long-tailed prompt tuning for image classification. In ICLR, 2022.
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2020.
  15. 15.Alexander Engelmann, Yuning Jiang, Boris Houska, and Timm Faulwasser. Decomposition of nonconvex optimization via bi-level distributed aladin. IEEE Transactions on Control of Network Systems, 2020.
  16. 16.Enrico Fini, Enver Sangineto, Stéphane Lathuilière, Zhun Zhong, Moin Nabi, and Elisa Ricci. A unified objective for novel class discovery. In ICCV, 2021.
  17. 17.Spyros Gidaris and Nikos Komodakis. Dynamic few-shot visual learning without forgetting. In CVPR, 2018.
  18. 18.Peiyan Gu, Chuyu Zhang, Ruijie Xu, and Xuming He. Class-relation knowledge distillation for novel class discovery. In ICCV, 2023.
  19. 19.Kai Han, Andrea Vedaldi, and Andrew Zisserman. Learning to discover novel visual categories via deep transfer clustering. In ICCV, 2019.
  20. 20.Kai Han, Sylvestre-Alvise Rebuffi, Sebastien Ehrhardt, Andrea Vedaldi, and Andrew Zisserman. Automatically discovering and learning new visual categories with ranking statistics. In ICLR, 2020.
  21. 21.Kai Han, Sylvestre-Alvise Rebuffi, Sebastien Ehrhardt, Andrea Vedaldi, and Andrew Zisserman. Autonovel: Automatically discovering and learning novel visual categories. IEEE TPAMI, 2021.
  22. 22.Shaozhe Hao, Kai Han, and Kwan-Yee K Wong. Cipr: An efficient framework with cross-instance positive relations for generalized category discovery. TMLR, 2024.
  23. 23.Jeff Z HaoChen, Colin Wei, Adrien Gaidon, and Tengyu Ma. Provable guarantees for self-supervised deep learning with spectral contrastive loss. In NeurIPS, 2021.
  24. 24.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  25. 25.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020.
  26. 26.Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. In ECCV, 2022.
  27. 27.Xuhui Jia, Kai Han, Yukun Zhu, and Bradley Green. Joint representation learning and novel category discovery on single- and multi-modal data. In ICCV, 2021.
  28. 28.Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. Maple: Multi-modal prompt learning. In CVPR, 2023.
  29. 29.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. In NeurIPS, 2020.
  30. 30.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In ICCV workshop, 2013.
  31. 31.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  32. 32.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Communications of the ACM, 2017.
  33. 33.Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi. Fine-grained visual classification of aircraft. arXiv preprint arXiv:1306.5151, 2013.
  34. 34.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  35. 35.Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski. DINOv2: Learning robust visual features without supervision. TMLR, 2024.
  36. 36.Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang. Moment matching for multi-source domain adaptation. In ICCV, 2019.
  37. 37.Nan Pu, Zhun Zhong, and Nicu Sebe. Dynamic conceptional contrastive learning for generalized category discovery. In CVPR, 2023.
  38. 38.Mamshad Nayeem Rizve, Navid Kardan, Salman Khan, Fahad Shahbaz Khan, and Mubarak Shah. Openldn: Learning to discover novel classes for open-world semi-supervised learning. In ECCV, 2022.
  39. 39.Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi. What does clip know about a red circle? visual prompt engineering for vlms. In ICCV, 2023.
  40. 40.Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. In NeurIPS, 2020.
  41. 41.Yiyou Sun, Zhenmei Shi, and Yixuan Li. A graph-theoretic framework for understanding open-world semi-supervised learning. In NeurIPS, 2024.
  42. 42.Kiat Chuan Tan, Yulong Liu, Barbara Ambrose, Melissa Tulig, and Serge Belongie. The herbarium challenge 2019 dataset. arXiv preprint arXiv: 1906.05372, 2019.
  43. 43.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In NeurIPS, 2017.
  44. 44.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive multiview coding. In ECCV, 2020.
  45. 45.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. JMLR, 2008.
  46. 46.Sagar Vaze, Kai Han, Andrea Vedaldi, and Andrew Zisserman. Open-set recognition: A good closed-set classifier is all you need. In ICLR, 2021.
  47. 47.Sagar Vaze, Kai Han, Andrea Vedaldi, and Andrew Zisserman. Generalized category discovery. In CVPR, 2022.
  48. 48.Sagar Vaze, Andrea Vedaldi, and Andrew Zisserman. No representation rules them all in category discovery. In NeurIPS, 2023.
  49. 49.Yabin Wang, Zhiwu Huang, and Xiaopeng Hong. S-prompts learning with pre-trained transformers: An occam’s razor for domain incremental learning. In NeurIPS, 2022a.
  50. 50.Yidong Wang, Hao Chen, Yue Fan, Wang Sun, Ran Tao, Wenxin Hou, Renjie Wang, Linyi Yang, Zhi Zhou, Lan-Zhe Guo, et al. Usb: A unified semi-supervised learning benchmark for classification. In NeurIPS, 2022b.
  51. 51.Yu Wang, Zhun Zhong, Pengchong Qiao, Xuxin Cheng, Xiawu Zheng, Chang Liu, Nicu Sebe, Rongrong Ji, and Jie Chen. Discover and align taxonomic context priors for open-world semi-supervised learning. In NeurIPS, 2024.
  52. 52.Peter Welinder, Steve Branson, Takeshi Mita, Catherine Wah, Florian Schroff, Serge Belongie, and Pietro Perona. Caltech-ucsd birds 200. 2010.
  53. 53.Xin Wen, Bingchen Zhao, and Xiaojuan Qi. Parametric classification for generalized category discovery: A baseline study. ICCV, 2023.
  54. 54.Sheng Zhang, Salman Khan, Zhiqiang Shen, Muzammal Naseer, Guangyi Chen, and Fahad Khan. Promptcal: Contrastive affinity learning via auxiliary prompts for generalized novel category discovery. In CVPR, 2023.
  55. 55.Bingchen Zhao and Kai Han. Novel visual category discovery with dual ranking statistics and mutual knowledge distillation. In NeurIPS, 2021.
  56. 56.Bingchen Zhao, Xin Wen, and Kai Han. Learning semi-supervised gaussian mixture models for generalized category discovery. In ICCV, 2023.
  57. 57.Zhun Zhong, Enrico Fini, Subhankar Roy, Zhiming Luo, Elisa Ricci, and Nicu Sebe. Neighborhood contrastive learning for novel class discovery. In CVPR, 2021a.
  58. 58.Zhun Zhong, Linchao Zhu, Zhiming Luo, Shaozi Li, Yi Yang, and Nicu Sebe. Openmix: Reviving known knowledge for discovering novel visual categories in an open world. In CVPR, 2021b.

Citation

MLA
Wang, H., et al. “SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning”. arXiv, 2024, http://arxiv.org/abs/2403.13684v3.
APA
Wang, H., Vaze, S., & Han, K. (2024). SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning. arXiv. http://arxiv.org/abs/2403.13684v3
Chicago
Wang, H., S. Vaze, and K. Han. 2024. “SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning”. arXiv. http://arxiv.org/abs/2403.13684v3.
Harvard
Wang, H., Vaze, S. and Han, K. (2024) “SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.13684v3.
Vancouver
1. Wang H, Vaze S, Han K (2024) SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning. arXiv

BibTeX

@article{wang2024sptnet,
  title = {SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning},
  author = {Wang, Hongjun and Vaze, Sagar and Han, Kai},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.13684v3},
  eprint = {2403.13684}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors