SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning
Hongjun WangSagar VazeKai Han
Introduces SPTNet, an efficient Generalized Category Discovery framework that alternates model fine-tuning with spatial prompt tuning to achieve a 10% accuracy gain over prior state-of-the-art baselines on the Semantic Shift Benchmark with only 0.117% parameter overhead.
Real-world computer vision systems frequently encounter open-world scenarios where incoming data includes both previously seen categories and completely novel, unseen categories. Generalized category discovery addresses this challenge by using partially labelled data from known categories to classify unlabelled images across both seen and novel classes. Standard approaches typically fine-tune large pre-trained vision models, which introduces significant computational costs and risks overfitting to the labelled data.
The main objective of the article is to demonstrate SPTNet, a highly parameter-efficient learning framework that integrates spatial prompt tuning to categorize unlabelled images across known and unknown classes without heavy model retraining.
The evaluated approach divides input images into local patches and injects learnable, pixel-level prompt parameters around each patch as well as a global border around the entire image. Instead of jointly optimizing the model and prompts—which leads to optimization instability—the framework alternates in two stages: freezing the vision backbone to update prompt parameters, and then freezing the prompts to fine-tune only the projection head and the top layer of the vision backbone using contrastive learning. This alternating optimization treats spatial prompts as targeted, learned data augmentations. The article evaluates this method across seven standard image benchmarks, spanning generic datasets (such as CIFAR-10, CIFAR-100, and ImageNet-100) and fine-grained benchmarks (such as CUB, Stanford Cars, FGVC-Aircraft, and Herbarium-19).
The article demonstrates several key findings:
- On the challenging fine-grained Semantic Shift Benchmark, the method achieves an average accuracy of 61.4%, surpassing prior state-of-the-art baselines by approximately 10% in proportional terms (and about 5% in absolute terms).
- Across generic image datasets, the approach consistently matches or exceeds existing methods, reaching 97.3% accuracy on CIFAR-10, 81.3% on CIFAR-100, and 85.4% on ImageNet-100.
- The full framework achieves these gains while introducing only 0.117% additional parameters relative to the base Vision Transformer architecture, with simpler patch-level variants requiring as little as 0.039% extra parameters.
- The alternating two-stage training scheme substantially outperforms standard end-to-end joint training, which causes prompts to degrade and become inactive, while cutting training time roughly in half compared to competing discovery methods.
These results show that adapting the input data representation through local spatial prompting is both more effective and far more resource-efficient than extensively retraining large backbone models. By directing the visual model's attention toward critical, fine-grained object parts, the framework enables strong knowledge transfer from known to unknown categories. This substantially reduces computing requirements and training turnaround times for deploying vision systems in dynamic environments.
Decision-makers and practitioners implementing open-world vision pipelines should adopt parameter-efficient spatial prompting and alternating optimization over full model fine-tuning. For operational deployment, smaller prompt sizes (such as a single-pixel border per patch) should be selected to avoid occluding core visual content. Future work should focus on validating the approach in production pilot environments and exploring more advanced foundation models like DINOv2 to further improve novel category discovery.
The framework exhibits certain limitations. Performance gains are less pronounced on low-resolution images (such as 32x32 pixels), where image patches lack sufficient spatial detail. Additionally, performance declines when models face significant cross-domain shifts or when the true category count is unknown, and the underlying decision-making of the learned prompts lacks full interpretability. Nonetheless, the consistent empirical improvements across multiple benchmarks provide high confidence in the framework's effectiveness for fine-grained open-world visual discovery.
- Paper: Generalized Category Discovery with Decoupled Prototypical Network, Wenbin An et al. (2023). This paper establishes the generalized category discovery formulation and prototype-based transfer methods that SPTNet directly aims to improve via efficient spatial prompt tuning.
- Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). This work introduces visual prompt tuning as a parameter-efficient alternative to full fine-tuning of vision Transformers, providing the conceptual foundation for SPTNet's spatial prompt design.
- Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). This study introduces context optimization and prompt tuning paradigms for adapting frozen foundation models, establishing the parameter-efficient adaptation principles leveraged by SPTNet.
- Paper: DINOv2: Learning Robust Visual Features without Supervision, Maxime Oquab et al. (2023). This paper provides key insights into self-supervised Vision Transformers and patch-level feature learning that underpin modern backbones and baseline representations evaluated in generalized category discovery.
No sufficiently relevant recommendations were found.
