Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models

Yabin ZhangWenjie ZhuHui TangZhiyuan MaKaiyang ZhouLei Zhang

article2024CVPR57 citations

Proposes Dual Memory Networks, a unified framework combining static training caches with dynamic test-time memory to adapt pre-trained vision-language models across zero-shot, few-shot, and training-free settings without relying on external data.

Listen

Pre-trained vision-language models like CLIP provide strong baseline capabilities for image recognition, yet adapting them to specific downstream classification tasks remains challenging. Existing adaptation techniques are largely fragmented: some require compute-intensive prompt tuning on small training datasets, others rely on generating synthetic data, and training-free methods typically operate under rigid constraints. Crucially, current approaches are designed for only one or two deployment scenarios and overlook the valuable information available within historical test data encountered during live inference.

The article introduces and evaluates Dual Memory Networks, a unified adaptation framework designed to operate effectively across three distinct paradigms: zero-shot adaptation (where no labeled training data exists), training-free few-shot adaptation, and standard few-shot adaptation (where limited training data is available for model updates). The core objective is to deliver state-of-the-art classification performance without requiring external synthetic data or expensive test-time optimization.

The approach introduces a dual-memory mechanism comprising a static memory and a dynamic memory. The static memory caches visual features from labeled training data when available, while the dynamic memory updates online during inference to store high-confidence visual features from previously seen test samples. Both components share a cross-attention interaction module that produces sample-adaptive visual classifiers alongside standard text classifiers. The framework can function in a strictly training-free mode using base model representations or be enhanced in few-shot settings by training lightweight projection layers. The authors evaluated this system across 11 benchmark image classification datasets and four out-of-distribution robustness benchmarks using ResNet-50 and Vision Transformer backbones.

The evaluation yielded three primary findings. First, in zero-shot adaptation, the proposed framework surpassed existing methods without external training data by 3.40% on ResNet-50 (reaching 63.71% average accuracy) and by 5.27% on Vision Transformers (reaching 70.72% average accuracy). It also outperformed more complex methods that rely on generating synthetic training images with external diffusion models, beating CaFo by 1.48%. Second, the method achieved leading accuracy across both training-free and fine-tuned few-shot settings, establishing consistent improvements on benchmark tasks such as ImageNet across various sample sizes. Third, the dynamic memory mechanism demonstrated superior robustness under natural distribution shifts, maintaining high accuracy when tested on corrupted and domain-shifted datasets without requiring task-specific fine-tuning.

These findings indicate that actively preserving historical test features during deployment is significantly more effective and computationally efficient than relying on synthetic data generation or test-time back-propagation. The architecture processes test samples in roughly 10.7 milliseconds, which is over 40 times faster than test-time optimization methods that require hundreds of milliseconds per sample. This balance of high accuracy and low latency lowers operational compute costs and reduces latency risks in real-time vision pipelines.

Organizations deploying vision-language models should consider dual-memory architectures when migrating models to new visual domains, particularly when labeled training data is scarce or latency budgets are tight. Decision-makers can adopt the training-free dynamic mode for instant zero-shot deployment or apply lightweight projection fine-tuning when small batches of labeled data and brief training windows are permissible.

A key operational limitation is the memory footprint required to maintain feature caches across many categories. For instance, on a 1,000-class benchmark like ImageNet, dynamic and static caches require approximately 204.8 MB and 65.5 MB of storage, respectively. This overhead makes the approach less suitable for strictly memory-constrained edge hardware. The reported performance gains are well-supported across standard academic vision benchmarks, though prospective users should validate memory retention dynamics in target deployment environments characterized by severe class imbalance or noisy test streams.

No sufficiently relevant recommendations were found.

Cover for Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models

Abstract

With the emergence of pre-trained vision-language models like CLIP, how to adapt them to various downstream classification tasks has garnered significant attention in recent research. The adaptation strategies can be typically categorized into three paradigms: zero-shot adaptation, few-shot adaptation, and the recently-proposed training-free few-shot adaptation. Most existing approaches are tailored for a specific setting and can only cater to one or two of these paradigms. In this paper, we introduce a versatile adaptation approach that can effectively work under all three settings. Specifically, we propose the dual memory networks that comprise dynamic and static memory components. The static memory caches training data knowledge, enabling training-free few-shot adaptation, while the dynamic memory preserves historical test features online during the testing process, allowing for the exploration of additional data insights beyond the training set. This novel capability enhances model performance in the few-shot setting and enables model usability in the absence of training data. The two memory networks employ the same flexible memory interactive strategy, which can operate in a training-free mode and can be further enhanced by incorporating learnable projection layers. Our approach is tested across 11 datasets under the three task settings. Remarkably, in the zero-shot scenario, it outperforms existing methods by over 3% and even shows superior results against methods utilizing external training data. Additionally, our method exhibits robust performance against natural distribution shifts. Codes are available at https://github.com/YBZh/DMN.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. A Flexible Memory Interactive Strategy
  • 3.2. Dynamic Memory Network
  • 3.3. Dual Memory Networks
  • 4. Experiments
  • 4.1. Experiment Settings
  • 4.2. Performance Evaluation
  • 4.3. Ablation and Analyses
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Dual Memory Networks Architecture for Vision-Language Adaptation

    model/method

    Dual Memory Networks (DMN) adapt pre-trained vision-language models (such as CLIP) to downstream classification tasks across zero-shot, few-shot, and training-free few-shot settings. DMN integrates knowledge from three distinct representations: the pre-trained text classifier, dynamic memory caching historical test features, and static memory caching few-shot training features.

    Given an image visual feature v∈RDv \in \mathbb{R}^D and category text embeddings C∈RC×DC \in \mathbb{R}^{C \times D} for CC downstream classes (L2L_2-normalized along feature dimension DD), the zero-shot prediction probability distribution from text is: Pt=Softmax(vC⊤)∈RCP^t = \text{Softmax}\left(v C^\top\right) \in \mathbb{R}^C

    DMN computes the final classification prediction probability vector Pdmn∈RCP^{dmn} \in \mathbb{R}^C by linearly ensembling the text prediction PtP^t, the dynamic memory prediction PdP^d, and the static memory prediction PsP^s: Pdmn=α1Pt+α2Pd+α3PsP^{dmn} = \alpha_1 P^t + \alpha_2 P^d + \alpha_3 P^s where α1,α2,α3≥0\alpha_1, \alpha_2, \alpha_3 \ge 0 are task-specific weighting hyperparameters.

    The framework configures three distinct task variants:

    • DMN-ZS (Zero-Shot): The dynamic memory MdM^d is activated online; static memory MsM^s is omitted (α3=0\alpha_3 = 0); memory projection layers are set to identity ( extLinear∗(x)=0\ ext{Linear}_*(x) = 0). No training data or test-time backpropagation is used.
    • DMN-TF (Training-Free Few-Shot): Both dynamic memory MdM^d and static memory MsM^s are active; projection layers remain frozen in identity mode, avoiding parameter training.
    • DMN (Few-Shot): Both MdM^d and MsM^s are active, and projection layers are optimized on labeled few-shot training data via gradient descent.
  2. Knowl 2 — Dynamic Memory Network and Online Test-Feature Caching

    model/method

    The dynamic memory network in Dual Memory Networks (DMN) caches features of historical test samples during inference to adaptively update classification prototypes on-the-fly without gradient backpropagation.

    1. Memory Structure and Initialization: The dynamic memory tensor is defined as Md∈RC×L×DM^d \in \mathbb{R}^{C \times L \times D}, initialized to zero, where CC is the number of target classes, LL is the memory capacity per category (set to L=50L=50 by default), and DD is the feature dimension. The extended dynamic memory is constructed by appending the text classifier embeddings: M~d=[Md,C]∈RC×(L+1)×D\widetilde{M}^d = [M^d, C] \in \mathbb{R}^{C \times (L+1) \times D}.
    2. Pseudo-Labeling and Memory Update: For each incoming test sample with feature v∈RDv \in \mathbb{R}^D, a pseudo-label y∈{1,…,C}y \in \{1, \dots, C\} is assigned via the zero-shot text prediction: y=arg⁡max⁡jPjty = \arg\max_j P^t_j. Feature vv is written into an empty row of sub-memory Myd∈RL×DM^d_y \in \mathbb{R}^{L \times D}. When MydM^d_y is full, the prediction entropy of vv (derived from PtP^t) is compared against stored samples; if vv has lower entropy than the stored sample with the highest entropy, it replaces that sample.
    3. Classifier Readout and Prediction: An adaptive classifier Cd∈RC×DC^d \in \mathbb{R}^{C \times D} is retrieved from extended dynamic memory M~d\widetilde{M}^d via cross-attention readout: Cd=ReadOut(v,M~d)C^d = \text{ReadOut}\left(v, \widetilde{M}^d\right). The dynamic classification probability is then computed as: Pd=Softmax(v(Cd)⊤)∈RCP^d = \text{Softmax}\left(v (C^d)^\top\right) \in \mathbb{R}^C
  3. Knowl 3 — Cross-Attention Memory Readout Mechanism

    equation

    Given a query test feature v∈RDv \in \mathbb{R}^D and a category-split memory tensor M∈RC×L×DM \in \mathbb{R}^{C \times L \times D} (where My∈RL×DM_y \in \mathbb{R}^{L \times D} represents the memory block for class yy), the sample-adaptive classifier Cm=ReadOut(v,M)∈RC×DC^m = \text{ReadOut}(v, M) \in \mathbb{R}^{C \times D} is defined for each row y∈{1,…,C}y \in \{1, \dots, C\} by: Cym=ωo(φ(ωq(v)ωk(My)⊤)ωv(My))C^m_y = \omega_o\left(\varphi\left(\omega_q(v) \omega_k(M_y)^\top\right) \omega_v(M_y)\right) where:

    • ωq,ωk,ωv,ωo:RD→RD\omega_q, \omega_k, \omega_v, \omega_o: \mathbb{R}^D \to \mathbb{R}^D are projection functions for query, key, value, and output respectively.
    • ωq(v)ωk(My)⊤∈R1×L\omega_q(v) \omega_k(M_y)^\top \in \mathbb{R}^{1 \times L} calculates cosine similarities between the projected query and the LL cached vectors of class yy.
    • φ(x):R→R\varphi(x): \mathbb{R} \to \mathbb{R} is an element-wise exponential modulation function that sharpens the similarity weights: φ(x)=exp⁡(−β(1−x))\varphi(x) = \exp(-\beta(1-x)) where β>0\beta > 0 is a sharpness hyperparameter (fixed to β=5.5\beta=5.5).
    • The memory-derived prediction probability vector Pm∈RCP^m \in \mathbb{R}^C is obtained by: Pm=Softmax(v(Cm)⊤)P^m = \text{Softmax}\left(v (C^m)^\top\right)
  4. Knowl 4 — Residual Feature Projection Layer for Memory Interaction

    model/method

    The projection functions ω∗∈{ωq,ωk,ωv,ωo}\omega_* \in \{\omega_q, \omega_k, \omega_v, \omega_o\} used in the memory readout mechanism are parameterized with a residual architecture: ω∗(x)=L2(x+Linear∗(x))\omega_*(x) = L_2\left(x + \text{Linear}_*(x)\right) where x∈RDx \in \mathbb{R}^D is an L2L_2-normalized input feature vector, Linear∗(x)\text{Linear}_*(x) denotes a linear layer initialized with weights and biases set to zero, and L2(⋅)L_2(\cdot) denotes L2L_2 normalization along the feature dimension.

    This parameterization operates in two modes:

    • Training-Free Mode (Zero-shot / DMN-TF): When all linear layers remain unoptimized (zero weights), ω∗(x)\omega_*(x) simplifies to the identity mapping ω∗(x)=x\omega_*(x) = x, allowing parameter-free memory interaction in the original CLIP embedding space.
    • Optimized Mode (Few-Shot DMN): When labeled training data are available, the linear layers are optimized with classification loss using AdamW, transforming feature embeddings into a more discriminative space for memory interaction.
  5. Knowl 5 — Static Memory Network for Few-Shot Training Data

    model/method

    In few-shot learning with CC classes and KK labeled examples per class (CC-way KK-shot), the static memory network stores visual features of all C×KC \times K training images in a fixed memory tensor Ms∈RC×K×DM^s \in \mathbb{R}^{C \times K \times D}.

    Unlike dynamic memory MdM^d, static memory MsM^s remains unaltered during evaluation after construction. Given a test feature v∈RDv \in \mathbb{R}^D, an adaptive static classifier Cs∈RC×DC^s \in \mathbb{R}^{C \times D} and its corresponding prediction probability Ps∈RCP^s \in \mathbb{R}^C are computed as: Cs=ReadOut(v,Ms)C^s = \text{ReadOut}\left(v, M^s\right) Ps=Softmax(v(Cs)⊤)P^s = \text{Softmax}\left(v (C^s)^\top\right)

    Maintaining a dedicated static memory prevents labeled training knowledge from being diluted as dynamic test memory fills up during online inference.

  6. Knowl 6 — Zero-Shot Classification Performance Across 11 Benchmarks

    data/table

    The zero-shot classification performance of DMN-ZS was evaluated on 11 downstream datasets (ImageNet, Flowers102, DTD, OxfordPets, StanfordCars, UCF101, Caltech101, Food101, SUN397, FGVCAircraft, EuroSAT) without external training data, using ResNet-50 and ViT-B/16 backbones.

    Method ImageNet Flower DTD Pets Cars UCF Caltech Food SUN Aircraft EuroSAT Mean
    ResNet-50 Backbone
    CLIP-RN50 58.16 61.75 40.37 83.57 55.70 58.84 85.88 73.97 58.80 15.66 23.69 56.04
    DN 60.16 63.32 41.21 81.92 56.55 55.60 87.25 74.64 59.11 17.43 28.31 56.86
    TPT 60.74 62.69 40.84 84.49 58.46 60.82 87.02 74.88 61.46 17.58 28.33 57.94
    VisDesc 59.68 65.37 41.96 82.39 54.76 58.47 88.11 76.80 59.84 16.26 37.60 58.29
    Ensemble 60.32 66.10 40.07 85.83 55.71 61.33 83.94 77.32 58.53 17.10 37.54 58.53
    CALIP 60.57 66.38 42.39 86.21 56.27 61.72 87.71 77.42 58.59 17.76 38.90 59.45
    DiffTPT∗^* 60.80 63.53 40.72 83.40 60.71 62.67 86.89 79.21 62.72 17.60 41.04 59.94
    CuPL 61.45 65.44 48.64 84.84 57.28 58.97 89.29 76.94 62.55 19.59 38.38 60.31
    SuS-X-SD-C∗^* 61.84 67.72 50.59 85.34 57.27 61.54 89.53 77.58 62.95 19.47 45.57 61.76
    CaFo∗^* 62.74 66.54 50.24 87.49 58.45 63.67 90.91 77.53 63.16 21.06 42.73 62.23
    DMN-ZS (Ours) 63.87 67.93 50.41 86.78 60.02 65.34 90.14 76.70 64.39 22.77 48.72 63.71
    ViT-B/16 Backbone
    CLIP-ViTB/16 66.73 67.44 44.27 88.25 65.48 65.13 93.35 83.65 62.59 23.67 42.01 63.87
    Ensemble 68.34 66.99 45.04 86.92 66.11 65.16 93.55 82.86 65.63 23.22 50.42 64.93
    TPT 68.98 68.98 47.75 87.79 66.87 68.04 94.16 84.67 65.50 24.78 42.44 65.45
    DiffTPT∗^* 70.30 70.10 47.00 88.20 67.01 68.22 92.49 87.23 65.74 25.60 43.13 65.91
    DMN-ZS (Ours) 72.25 74.49 55.85 92.04 67.96 72.51 95.38 85.08 70.18 30.03 59.43 70.72

    Methods labeled with ∗^* employ synthetic training data generated via diffusion or retrieval models. DMN-ZS achieves a mean accuracy of 63.71% (ResNet-50) and 70.72% (ViT-B/16), outperforming non-external-data methods by 4.26% and 5.27% over baseline CLIP/TPT, and exceeding external-data-augmented methods such as CaFo (62.23%) and DiffTPT (65.91%).

  7. Knowl 7 — Robustness to Natural Distribution Shifts on ImageNet Variants

    data/table

    The robustness of DMN-ZS was evaluated on standard ImageNet and four distribution-shifted benchmarks: ImageNet-A (-A), ImageNet-V2 (-V2), ImageNet-R (-R), and ImageNet-Sketch (-Sketch).

    Method ImageNet -A -V2 -R -Sketch
    ResNet-50 Backbone
    CLIP-RN50 58.16 21.83 51.41 56.15 33.37
    Ensemble 59.81 23.24 52.91 60.72 35.48
    TPT 60.74 26.67 54.70 59.11 35.09
    CALIP 60.57 23.96 53.70 60.81 35.61
    DiffTPT 60.80 31.06 55.80 58.80 37.10
    CoCoOp∗^* 62.81 23.32 55.72 57.74 34.48
    CoOp∗^* 63.33 23.06 55.40 56.60 34.67
    DMN-ZS (Ours) 63.87 28.57 56.12 61.44 39.84
    ViT-B/16 Backbone
    CLIP-ViT-B/16 66.73 47.87 60.86 73.98 46.09
    Ensemble 68.34 49.89 61.88 77.65 48.24
    TPT 68.98 54.77 63.45 77.06 47.94
    DiffTPT 70.30 55.68 65.10 75.00 46.80
    MaPLe∗^* 70.72 50.90 64.07 76.98 49.15
    CoCoOp∗^* 71.02 50.63 64.07 76.18 48.75
    CoOp∗^* 71.51 49.71 64.20 75.21 47.99
    PromptSRC∗^* 71.27 50.90 64.35 77.80 49.55
    DMN-ZS (Ours) 72.25 58.28 65.17 78.55 53.20

    Methods marked with ∗^* require 16-shot labeled ImageNet training data, while other methods require zero labeled training data. DMN-ZS outperforms both zero-shot and 16-shot tuned adaptation models across all out-of-distribution ImageNet evaluation sets on both backbones.

  8. Knowl 8 — Computational and Inference Efficiency Comparison

    data/table

    Computational efficiency was evaluated on zero-shot and 16-shot ImageNet adaptation using a ResNet-50 visual encoder on an NVIDIA RTX A6000 GPU.

    Methods Training Time Test Latency GFLOPs Learnable Parameters
    Zero-shot
    CLIP – 10.1 ms 0 0
    CALIP – 10.2 ms 0 0
    TPT – 436.0 ms >10>10 0.01 M
    DMN-ZS (Ours) – 10.7 ms 0 0
    Few-shot (16-shot)
    Tip-Adapter – 10.4 ms 0 0
    APE – 10.4 ms 0 0
    DMN-TF (Ours) – 10.7 ms 0 0
    CoOp 14 h 10.2 ms >10>10 0.01 M
    CLIP-Adapter 50 min 10.4 ms 0.004 0.52 M
    Tip-Adapter-F 5 min 10.4 ms 0.030 16.3 M
    APE-T 5 min 10.4 ms 0.002 0.51 M
    DMN (Ours) 5 min 10.7 ms 0.033 4.20 M

    DMN-ZS and DMN-TF execute in training-free mode with 0 GFLOPs and 0 learnable parameters, adding only 0.6 ms of test latency over standard CLIP (10.7 ms vs. 10.1 ms) and avoiding the 436 ms test-time gradient optimization overhead of TPT. In the trainable few-shot setting, DMN trains in 5 minutes with 4.20M parameters.

  9. Knowl 9 — Ablation on Dynamic Memory, Static Memory, Buffer Length, and Sharpness

    empirical result

    Ablation studies on ImageNet establish the contributions of individual DMN components:

    1. Dynamic vs. Static Memory: When tested separately in training-free few-shot classification, dynamic memory alone outperforms static memory alone across all shot counts (1, 2, 4, 8, 16 shots). Combining dynamic and static memories yields strictly superior performance to either single-memory configuration across all shots.
    2. Dynamic Memory Length (LL): In zero-shot classification (DMN-ZS), classification accuracy increases with memory capacity per class LL, growing rapidly from L=1L=1 to L=30L=30 and plateauing around L=50L=50.
    3. Projection Layer Configuration (Q,K,V,OQ, K, V, O): In few-shot DMN, adding learnable residual projection layers ωq,ωk,ωv,ωo\omega_q, \omega_k, \omega_v, \omega_o (denoted Q,K,V,OQ, K, V, O) progressively increases accuracy. The output projection ωo\omega_o provides the largest individual performance boost, and optimizing all four projections (QKVOQKVO) delivers the highest overall accuracy.
    4. Sharpness Hyperparameter (β\beta): The sharpness modulation parameter φ(x)=exp⁡(−β(1−x))\varphi(x) = \exp(-\beta(1-x)) achieves peak accuracy at β=5.5\beta = 5.5, with accuracy degrading when β\beta is set below 3.5 or above 7.5.
  10. Knowl 10 — Storage Overhead Limitation of Dual Memory Components

    limitation

    Maintaining explicit dynamic and static feature memory caches creates storage overhead that scales with the number of downstream classes CC, memory length LL, shot number KK, and feature dimension DD.

    For 16-shot ImageNet adaptation with C=1000C = 1000 classes and CLIP feature dimension D=1024D = 1024:

    • The dynamic memory Md∈RC×L×DM^d \in \mathbb{R}^{C \times L \times D} with L=50L=50 requires storing 5×1075 \times 10^7 floating-point values, occupying approximately 204.8 MB204.8\,\text{MB} of memory.
    • The static memory Ms∈RC×K×DM^s \in \mathbb{R}^{C \times K \times D} with K=16K=16 requires storing 1.6×1071.6 \times 10^7 floating-point values, occupying approximately 65.5 MB65.5\,\text{MB} of memory.

    This combined memory footprint of over 270 MB270\,\text{MB} may pose constraints when deploying to memory-restricted edge devices or scaling to classification tasks with very large label spaces.

Coverage note — None was omitted; all contributed models, equations, variants, empirical results across 11 datasets and distribution shift benchmarks, efficiency measurements, ablations, and limitations were fully covered.

References

  1. 1.Alan Baddeley. The episodic buffer: a new component of working memory? Trends in cognitive sciences, 4(11):417–423, 2000.
  2. 2.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101–mining discriminative components with random forests. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VI 13, pages 446–461. Springer, 2014.
  3. 3.Guangyi Chen, Weiran Yao, Xiangchen Song, Xinyue Li, Yongming Rao, and Kun Zhang. Prompt learning with optimal transport for vision-language models. arXiv preprint arXiv:2210.01253, 2022.
  4. 4.Yihong Chen, Yue Cao, Han Hu, and Liwei Wang. Memory enhanced global-local aggregation for video object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10337–10346, 2020.
  5. 5.Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3606–3613, 2014.
  6. 6.Katherine Crowson, Stella Biderman, Daniel Kornis, Dashiell Stander, Eric Hallahan, Louis Castricato, and Edward Raff. Vqgan-clip: Open domain image generation and editing with natural language guidance. In European Conference on Computer Vision, pages 88–105. Springer, 2022.
  7. 7.Boris Dayma, Suraj Patil, Pedro Cuenca, Khalid Saifullah, Tanishq Abraham, Phuc Le Khac, Luke Melas, and Ritobrata Ghosh. Dall· e mini. HuggingFace. com. https://huggingface. co/spaces/dallemini/dallemini (accessed Sep. 29, 2022), 2021.
  8. 8.Hanming Deng, Yang Hua, Tao Song, Zongpu Zhang, Zhengui Xue, Ruhui Ma, Neil Robertson, and Haibing Guan. Object guided external memory network for video object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6678–6687, 2019.
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  11. 11.Li Fei-Fei, Rob Fergus, and Pietro Perona. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In 2004 conference on computer vision and pattern recognition workshop, pages 178–178. IEEE, 2004.
  12. 12.Chun-Mei Feng, Kai Yu, Yong Liu, Salman Khan, and Wangmeng Zuo. Diverse data augmentation with diffusions for effective test-time prompt tuning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2704–2714, 2023.
  13. 13.Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. Clip-adapter: Better vision-language models with feature adapters. International Journal of Computer Vision, pages 1–15, 2023.
  14. 14.Ziyu Guo, Renrui Zhang, Longtian Qiu, Xianzheng Ma, Xupeng Miao, Xuming He, and Bin Cui. Calip: Zero-shot enhancement of clip with parameter-free attention. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 746–754, 2023.
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  16. 16.Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7):2217–2226, 2019.
  17. 17.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. The many faces of robustness: A critical analysis of out-of-distribution generalization. ICCV, 2021.
  18. 18.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. CVPR, 2021.
  19. 19.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pages 2790–2799. PMLR, 2019.
  20. 20.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning, pages 4904–4916. PMLR, 2021.
  21. 21.Geethan Karunaratne, Manuel Schmuck, Manuel Le Gallo, Giovanni Cherubini, Luca Benini, Abu Sebastian, and Abbas Rahimi. Robust high-dimensional memory-augmented neural networks. Nature communications, 12(1):2468, 2021.
  22. 22.Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. Maple: Multi-modal prompt learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19113–19122, 2023.
  23. 23.Muhammad Uzair Khattak, Syed Talal Wasim, Muzammal Naseer, Salman Khan, Ming-Hsuan Yang, and Fahad Shahbaz Khan. Self-regulating prompts: Foundational model adaptation without forgetting. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15190–15200, 2023.
  24. 24.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4015–4026, 2023.
  25. 25.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops, pages 554–561, 2013.
  26. 26.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021.
  27. 27.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems, 34:9694–9705, 2021.
  28. 28.Minghan Li, Shuai Li, Wangmeng Xiang, and Lei Zhang. Mdqe: Mining discriminative query embeddings to segment occluded instances on challenging videos. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10524–10533, 2023.
  29. 29.Minghan Li, Shuai Li, Xindong Zhang, and Lei Zhang. Univs: Unified and universal video segmentation with prompts as queries. arXiv preprint arXiv:2402.18115, 2024.
  30. 30.Shuai Li, Chenhang He, Ruihuang Li, and Lei Zhang. A dual weighting label assignment scheme for object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9387–9396, 2022.
  31. 31.Shuai Li, Minghan Li, Ruihuang Li, Chenhang He, and Lei Zhang. One-to-few label assignment for end-to-end dense detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7350–7359, 2023.
  32. 32.Shuai Li, Minghan Li, Pengfei Wang, and Lei Zhang. Opensd: Unified open-vocabulary segmentation and detection. arXiv preprint arXiv:2312.06703, 2023.
  33. 33.Jiehong Lin, Lihua Liu, Dekun Lu, and Kui Jia. Sam-6d: Segment anything model meets zero-shot 6d object pose estimation. arXiv preprint arXiv:2311.15707, 2023.
  34. 34.Zhiqiu Lin, Samuel Yu, Zhiyi Kuang, Deepak Pathak, and Deva Ramanan. Multimodality helps unimodality: Crossmodal few-shot learning with multimodal models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19325–19337, 2023.
  35. 35.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  36. 36.Yuning Lu, Jianzhuang Liu, Yonggang Zhang, Yajing Liu, and Xinmei Tian. Prompt distribution learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5206–5215, 2022.
  37. 37.Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi. Fine-grained visual classification of aircraft. arXiv preprint arXiv:1306.5151, 2013.
  38. 38.Sachit Menon and Carl Vondrick. Visual classification via description from large language models. arXiv preprint arXiv:2210.07183, 2022.
  39. 39.Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. In 2008 Sixth Indian conference on computer vision, graphics & image processing, pages 722–729. IEEE, 2008.
  40. 40.Zachary Novack, Julian McAuley, Zachary Chase Lipton, and Saurabh Garg. Chils: Zero-shot image classification with hierarchical label sets. In International Conference on Machine Learning, pages 26342–26362. PMLR, 2023.
  41. 41.Seoung Wug Oh, Joon-Young Lee, Ning Xu, and Seon Joo Kim. Video object segmentation using space-time memory networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9226–9235, 2019.
  42. 42.Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar. Cats and dogs. In 2012 IEEE conference on computer vision and pattern recognition, pages 3498–3505. IEEE, 2012.
  43. 43.Sarah Pratt, Ian Covert, Rosanne Liu, and Ali Farhadi. What does a platypus look like? generating customized prompts for zero-shot image classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15691–15701, 2023.
  44. 44.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  45. 45.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pages 5389–5400. PMLR, 2019.
  46. 46.Zhiyuan Ren, Yiyang Su, and xiaoming Liu. Chatgptpowered hierarchical comparisons for image classification. Advances in neural information processing systems, 2023.
  47. 47.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022.
  48. 48.Aditya Sanghi, Rao Fu, Vivian Liu, Karl DD Willis, Hooman Shayani, Amir H Khasahmadi, Srinath Sridhar, and Daniel Ritchie. Clip-sculptor: Zero-shot generation of high-fidelity and diverse shapes from natural language. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18339–18348, 2023.
  49. 49.Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap. Meta-learning with memory-augmented neural networks. In International conference on machine learning, pages 1842–1850. PMLR, 2016.
  50. 50.Cheng Shi and Sibei Yang. Logoprompt:synthetic text images can be good visual prompts for vision-language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023.
  51. 51.Manli Shu, Weili Nie, De-An Huang, Zhiding Yu, Tom Goldstein, Anima Anandkumar, and Chaowei Xiao. Testtime prompt tuning for zero-shot generalization in visionlanguage models. Advances in Neural Information Processing Systems, 35:14274–14289, 2022.
  52. 52.Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012.
  53. 53.Mark G Stokes. ‘activity-silent’working memory in prefrontal cortex: a dynamic coding framework. Trends in cognitive sciences, 19(7):394–405, 2015.
  54. 54.Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al. Endto-end memory networks. Advances in neural information processing systems, 28, 2015.
  55. 55.Lingchen Sun, Rongyuan Wu, Zhengqiang Zhang, Hongwei Yong, and Lei Zhang. Improving the stability of diffusion models for content consistent super-resolution. arXiv preprint arXiv:2401.00877, 2023.
  56. 56.Vishaal Udandarao, Ankush Gupta, and Samuel Albanie. Sus-x: Training-free name-only transfer of vision-language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2725–2736, 2023.
  57. 57.Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing. Learning robust global representations by penalizing local predictive power. Advances in Neural Information Processing Systems, 32, 2019.
  58. 58.Jason Weston, Sumit Chopra, and Antoine Bordes. Memory networks. arXiv preprint arXiv:1410.3916, 2014.
  59. 59.Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, et al. Robust fine-tuning of zero-shot models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7959–7971, 2022.
  60. 60.Rongyuan Wu, Tao Yang, Lingchen Sun, Zhengqiang Zhang, Shuai Li, and Lei Zhang. Seesr: Towards semanticsaware real-world image super-resolution. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024.
  61. 61.Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba. Sun database: Large-scale scene recognition from abbey to zoo. In 2010 IEEE computer society conference on computer vision and pattern recognition, pages 3485–3492. IEEE, 2010.
  62. 62.Guo-Sen Xie, Huan Xiong, Jie Liu, Yazhou Yao, and Ling Shao. Few-shot semantic segmentation with cyclic memory network. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7293–7302, 2021.
  63. 63.Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang, Guoqiang Liang, Peng Wang, and Yanning Zhang. Dual modality prompt tuning for vision-language pre-trained model. IEEE Transactions on Multimedia, 2023.
  64. 64.Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi, and Junzhou Huang. Vision-language pre-training with triple contrastive learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15671–15680, 2022.
  65. 65.Tao Yu, Zhihe Lu, Xin Jin, Zhibo Chen, and Xinchao Wang. Task residual for tuning vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10899–10909, 2023.
  66. 66.Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, and Chen Change Loy. Unified vision and language prompt learning. arXiv preprint arXiv:2210.07225, 2022.
  67. 67.Haojie Zhang, Yongyi Su, Xun Xu, and Kui Jia. Improving the generalization of segmentation foundation model under distribution shift via weakly supervised adaptation. arXiv preprint arXiv:2312.03502, 2023.
  68. 68.Renrui Zhang, Rongyao Fang, Wei Zhang, Peng Gao, Kunchang Li, Jifeng Dai, Yu Qiao, and Hongsheng Li. Tip-adapter: Training-free clip-adapter for better visionlanguage modeling. arXiv preprint arXiv:2111.03930, 2021.
  69. 69.Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li. Pointclip: Point cloud understanding by clip. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8552–8562, 2022.
  70. 70.Renrui Zhang, Xiangfei Hu, Bohao Li, Siyuan Huang, Hanqiu Deng, Yu Qiao, Peng Gao, and Hongsheng Li. Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15211–15222, 2023.
  71. 71.Yabin Zhang, Bin Deng, Hui Tang, Lei Zhang, and Kui Jia. Unsupervised multi-class domain adaptation: Theory, algorithms, and practice. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(5):2775–2792, 2020.
  72. 72.Yabin Zhang, Minghan Li, Ruihuang Li, Kui Jia, and Lei Zhang. Exact feature distribution matching for arbitrary style transfer and domain generalization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8035–8045, 2022.
  73. 73.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16816–16825, 2022.
  74. 74.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision, 130(9):2337–2348, 2022.
  75. 75.Yifei Zhou, Juntao Ren, Fengyu Li, Ramin Zabih, and Ser-Nam Lim. Distribution normalization: An” effortless” testtime augmentation for contrastively learned visual-language models. arXiv preprint arXiv:2302.11084, 2023.
  76. 76.Beier Zhu, Yulei Niu, Yucheng Han, Yue Wu, and Hanwang Zhang. Prompt-aligned gradient for prompt tuning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15659–15669, 2023.
  77. 77.Xiangyang Zhu, Renrui Zhang, Bowei He, Aojun Zhou, Dong Wang, Bin Zhao, and Peng Gao. Not all features matter: Enhancing few-shot clip with adaptive prior refinement. arXiv preprint arXiv:2304.01195, 2023.

Citation

MLA
Zhang, Y., et al. “Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models”. arXiv, 2024, http://arxiv.org/abs/2403.17589v1.
APA
Zhang, Y., Zhu, W., Tang, H., Ma, Z., Zhou, K., & Zhang, L. (2024). Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models. arXiv. http://arxiv.org/abs/2403.17589v1
Chicago
Zhang, Y., W. Zhu, H. Tang, Z. Ma, K. Zhou, and L. Zhang. 2024. “Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models”. arXiv. http://arxiv.org/abs/2403.17589v1.
Harvard
Zhang, Y. et al. (2024) “Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.17589v1.
Vancouver
1. Zhang Y, Zhu W, Tang H, Ma Z, Zhou K, Zhang L (2024) Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models. arXiv

BibTeX

@article{zhang2024dual,
  title = {Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models},
  author = {Zhang, Yabin and Zhu, Wenjie and Tang, Hui and Ma, Zhiyuan and Zhou, Kaiyang and Zhang, Lei},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.17589v1},
  eprint = {2403.17589}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE