Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
Yabin ZhangWenjie ZhuHui TangZhiyuan MaKaiyang ZhouLei Zhang
Proposes Dual Memory Networks, a unified framework combining static training caches with dynamic test-time memory to adapt pre-trained vision-language models across zero-shot, few-shot, and training-free settings without relying on external data.
Pre-trained vision-language models like CLIP provide strong baseline capabilities for image recognition, yet adapting them to specific downstream classification tasks remains challenging. Existing adaptation techniques are largely fragmented: some require compute-intensive prompt tuning on small training datasets, others rely on generating synthetic data, and training-free methods typically operate under rigid constraints. Crucially, current approaches are designed for only one or two deployment scenarios and overlook the valuable information available within historical test data encountered during live inference.
The article introduces and evaluates Dual Memory Networks, a unified adaptation framework designed to operate effectively across three distinct paradigms: zero-shot adaptation (where no labeled training data exists), training-free few-shot adaptation, and standard few-shot adaptation (where limited training data is available for model updates). The core objective is to deliver state-of-the-art classification performance without requiring external synthetic data or expensive test-time optimization.
The approach introduces a dual-memory mechanism comprising a static memory and a dynamic memory. The static memory caches visual features from labeled training data when available, while the dynamic memory updates online during inference to store high-confidence visual features from previously seen test samples. Both components share a cross-attention interaction module that produces sample-adaptive visual classifiers alongside standard text classifiers. The framework can function in a strictly training-free mode using base model representations or be enhanced in few-shot settings by training lightweight projection layers. The authors evaluated this system across 11 benchmark image classification datasets and four out-of-distribution robustness benchmarks using ResNet-50 and Vision Transformer backbones.
The evaluation yielded three primary findings. First, in zero-shot adaptation, the proposed framework surpassed existing methods without external training data by 3.40% on ResNet-50 (reaching 63.71% average accuracy) and by 5.27% on Vision Transformers (reaching 70.72% average accuracy). It also outperformed more complex methods that rely on generating synthetic training images with external diffusion models, beating CaFo by 1.48%. Second, the method achieved leading accuracy across both training-free and fine-tuned few-shot settings, establishing consistent improvements on benchmark tasks such as ImageNet across various sample sizes. Third, the dynamic memory mechanism demonstrated superior robustness under natural distribution shifts, maintaining high accuracy when tested on corrupted and domain-shifted datasets without requiring task-specific fine-tuning.
These findings indicate that actively preserving historical test features during deployment is significantly more effective and computationally efficient than relying on synthetic data generation or test-time back-propagation. The architecture processes test samples in roughly 10.7 milliseconds, which is over 40 times faster than test-time optimization methods that require hundreds of milliseconds per sample. This balance of high accuracy and low latency lowers operational compute costs and reduces latency risks in real-time vision pipelines.
Organizations deploying vision-language models should consider dual-memory architectures when migrating models to new visual domains, particularly when labeled training data is scarce or latency budgets are tight. Decision-makers can adopt the training-free dynamic mode for instant zero-shot deployment or apply lightweight projection fine-tuning when small batches of labeled data and brief training windows are permissible.
A key operational limitation is the memory footprint required to maintain feature caches across many categories. For instance, on a 1,000-class benchmark like ImageNet, dynamic and static caches require approximately 204.8 MB and 65.5 MB of storage, respectively. This overhead makes the approach less suitable for strictly memory-constrained edge hardware. The reported performance gains are well-supported across standard academic vision benchmarks, though prospective users should validate memory retention dynamics in target deployment environments characterized by severe class imbalance or noisy test streams.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. This foundational work introduces Contrastive Language-Image Pre-training (CLIP), providing the core zero-shot vision-language model architecture and representations that Dual Memory Networks adapt.
- Paper: CLIP-Adapter: Better Vision-Language Models with Feature Adapters, Peng Gao et al. (2021). It introduces parameter-efficient feature adapters with residual connections for adapting pre-trained CLIP models, establishing the lightweight adapter paradigm contrasted and refined in dual-memory adaptation.
- Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). It pioneers continuous prompt learning for vision-language models via Context Optimization (CoOp), framing the few-shot adaptation baseline that the source paper seeks to outperform with cache-based memory mechanisms.
- Paper: Conditional Prompt Learning for Vision-Language Models, Kaiyang Zhou et al. (2022). It demonstrates image-conditional prompt adaptation for vision-language models, addressing generalization and overfitting issues that motivate sample-adaptive classification.
- Paper: Feature Alignment and Uniformity for Test Time Adaptation, Shuai Wang et al. (2023). It establishes online test-time adaptation using memory banks and dynamic filtering on unlabeled inference streams, formulating the conceptual basis for dynamic memory caching.
- Paper: SwapPrompt: Test-Time Prompt Adaptation for Vision-Language Models, Xiaosong Ma et al. (2023). It develops test-time adaptation for vision-language models via dynamic and historical prompt updating, highlighting the computational trade-offs of test-time optimization that the source paper resolves.
- Paper: Meta-Learning with Memory-Augmented Neural Networks, Adam Santoro et al. (2016). It formalizes memory-augmented neural networks for rapid few-shot meta-learning and external slot-based information caching.
No sufficiently relevant recommendations were found.
