Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models

Didi ZhuZhongyi SunZexi LiTao ShenKe YanShouhong DingChao WuKun Kuang

article2024ICML50 citations

Proposes Model Tailor, a parameter-efficient post-training method that updates fewer than ten percent of model parameters through sparse masking and Hessian-based compensation, preventing catastrophic forgetting in multi-modal large language models while maintaining performance on both original and target tasks.

Listen

Adapting multimodal artificial intelligence models to new downstream applications often introduces a critical operational challenge known as catastrophic forgetting. When multimodal large language models—which integrate visual perception with complex language reasoning—are fine-tuned on new tasks such as image captioning or visual question answering, their ability to perform their original, foundational tasks frequently degrades by a substantial margin. Existing remedies designed for smaller models or text-only systems fail to address the complex cross-modal architecture and computational demands of large multimodal systems. Full retraining across all tasks remains computationally expensive, while standard parameter-efficient methods such as Low-Rank Adaptation continue to retain redundant parameter updates that overwrite pre-existing capabilities.

The article demonstrates and evaluates a post-training framework named Model Tailor, which aims to efficiently adapt multimodal models to target tasks while preserving their baseline proficiency. The core objective is to integrate task-specific modifications into pre-trained models by selectively replacing no more than 10% of the model parameters and mathematically compensating the retained weights to maintain high performance across both old and new domains.

The researchers evaluated this framework on two standard multimodal systems: InstructBLIP, focusing on adapting its 288-million-parameter vision-language connector, and LLaVA-1.5, focusing on adapting both its connector and deep language layers totaling 2.7 billion parameters. The method was tested across a diverse benchmark suite, fine-tuning on datasets such as Flickr30k, GQA, and OKVQA, and assessing performance retention across broad evaluation sets including COCO, VQAv2, VizWiz, and MM-Bench. The approach breaks down the optimization into layer-wise sub-tasks using second-order approximations, selecting high-impact parameters via a hybrid scoring mechanism that balances parameter change and task sensitivity, and adjusting retained weights using an inverse Hessian compensation method.

The findings confirm that Model Tailor preserves original task capabilities while successfully acquiring new skills. Across standard single-task benchmarks, the framework maintained approximately 99% of pre-trained baseline effectiveness while reaching roughly 97% of the performance achieved by standard full fine-tuning on the new target tasks. By contrast, conventional fine-tuning caused severe degradation; for example, LLaVA's accuracy on the VizWiz benchmark dropped from 50.0% to 27.24% following fine-tuning on Flickr30k. When tested in multi-task scenarios, the framework smoothly combined distinct task adjustments without suffering from performance drops across alternating objectives. Furthermore, when combined with Low-Rank Adaptation on LLaVA-1.5, Model Tailor achieved an 18.9-point gain on Flickr30k and improved the overall balanced evaluation score by 9.9%, demonstrating that it effectively removes extraneous parameter updates.

These results show that engineering teams can reliably specialize multimodal models for domain-specific deployments without sacrificing general intelligence or maintaining disconnected models for separate workflows. By modifying only a small fraction of parameters in a rapid post-training step, organizations can reduce the computing costs, infrastructure requirements, and deployment risks associated with maintaining separate large models. The post-training process requires minimal extra computation—taking roughly 21 minutes on a single graphics processing unit during testing—making it an operationally practical addition to standard training pipelines.

Organizations deploying multimodal foundation models should consider adopting sparse parameter selection and compensation strategies like Model Tailor over standard unconstrained fine-tuning. Engineering teams should also combine this post-training adjustment with parameter-efficient techniques such as Low-Rank Adaptation to prune redundant updates. While the framework relies on layer-wise approximations and a recommended sparsity budget near 10%, its strong performance across both single-task and multi-task settings offers high confidence for integrating specialized visual and reasoning capabilities into existing large model deployments.

  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Learning without Forgetting establishes the core stability–plasticity problem and a foundational approach to preserving old capabilities while fine-tuning for new tasks.
  • Paper: Memory Aware Synapses: Learning what (not) to forget, Rahaf Aljundi et al. (2017). Memory Aware Synapses introduces parameter-importance scoring to protect knowledge during updates, a key conceptual precursor to Model Tailor’s selective parameter retention.
  • Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). This unified account of parameter-efficient transfer learning clarifies methods such as LoRA that Model Tailor evaluates and combines with its sparse update strategy.
  • Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). This long-horizon study extends forgetting mitigation from multimodal post-training to persistent language-model learning, testing how protective mechanisms compose across many sequential updates.
  • Paper: Self-Distillation Enables Continual Learning, Idan Shenfeld et al. (2026). Self-Distillation Fine-Tuning continues the effort to preserve prior capabilities during adaptation, applying a teacher–student learning process to continual skill and knowledge acquisition.
Cover for Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models

Abstract

Catastrophic forgetting emerges as a critical challenge when fine-tuning multi-modal large language models (MLLMs), where improving performance on target tasks often leads to a significant performance drop on the original tasks. This paper presents a comprehensive analysis of catastrophic forgetting in MLLMs and introduces a post-training adjustment method called Model Tailor. Our method primarily preserves the pre-trained parameters while replacing a small number (≤ 10%) of fine-tuned parameters, maintaining ~ 99% effectiveness on original tasks versus pre-training, and achieving ~ 97% on new tasks compared to standard fine-tuning. Specifically, we derive a sparse mask to identify the “model patch”, based on a fusion strategy that integrates salience and sensitivity analysis. Subsequently, a compensation mechanism is introduced to “decorate the patch”, enhancing the model’s performance on both target and original tasks. Additionally, our method is adaptable to multi-task scenarios. Through extensive experiments on Instruct-BLIP and LLaVA-1.5 in both image captioning and visual question answering tasks, our approach demonstrates significant task adaptability while preserving inherent pre-trained capabilities.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Problem Formulation
  • 4. Method
  • 4.1. Overview of Model Tailor
  • 4.2. Layer-wise Objective of Model Tailor
  • 4.3. Identify Model Patch
  • 4.4. Decorate the Patch
  • 5. Experiments
  • 5.1. Experimental Setup
  • 5.2. Main Results on Single-Task setting
  • 5.3. Main Results on Multi-Task Setting
  • 5.4. Synergy with Parameter-Efficient Method
  • 5.5. Ablation Study
  • 5.6. Computation Cost
  • 6. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • Appendix of Model Tailor: Mitigating Catastrophic Forgetting in MLLMs
  • A. More Details about Related Work
  • B. More Details about Method
  • B.1. Preliminary
  • B.2. Extension to Multi-Task Scenario
  • C. Proof
  • C.1. Proof of Theorem 4.5
  • C.2. Proof of Theorem 4.3
  • D. More Details about Experiments
  • D.1. Implementation Details
  • D.2. Evaluation Metrics
  • D.3. More Results

Knowls

  1. Knowl 1 — Model Tailor fuses a sparse fine-tuned patch into the pretrained model

    model/method

    Let Θpre\Theta_{\mathrm{pre}} and Θsft\Theta_{\mathrm{sft}} be the pretrained and target-task fine-tuned parameters of a multimodal large language model (MLLM). Model Tailor constructs fused parameters as Θfusion=M⊙(Θsft+C)+(I−M)⊙Θpre\Theta_{\mathrm{fusion}}=M\odot(\Theta_{\mathrm{sft}}+C)+(I-M)\odot\Theta_{\mathrm{pre}}, where MM is a binary mask selecting the fine-tuned parameters retained as the model patch, CC contains compensation adjustments on selected parameters, II is the identity mask, and ⊙\odot denotes elementwise multiplication. Thus unselected parameters are restored to their pretrained values, while selected parameters use their fine-tuned values with compensation. The goal is to improve performance on a new target task TT while retaining performance on the original task set PP: the target loss increase relative to full fine-tuning, LT(Θfusion)−LT(Θsft)L_T(\Theta_{\mathrm{fusion}})-L_T(\Theta_{\mathrm{sft}}), should be at most ϵt\epsilon_t, and the original-task loss increase relative to pretraining, LP(Θfusion)−LP(Θpre)L_P(\Theta_{\mathrm{fusion}})-L_P(\Theta_{\mathrm{pre}}), should be at most ϵp\epsilon_p. The method is a post-training adjustment; it does not require retraining the full model.

  2. Knowl 2 — Layer-wise optimization approximates preservation without original training data

    model/method

    For a model layer ℓ\ell, let Θpreℓ\Theta^\ell_{\mathrm{pre}}, Θsftℓ\Theta^\ell_{\mathrm{sft}}, and Θfusionℓ\Theta^\ell_{\mathrm{fusion}} denote its pretrained, fine-tuned, and fused weights; let fℓ(X,Θ)f_\ell(X,\Theta) be its output on layer inputs XX. Model Tailor frames fusion as matching the fine-tuned layer output on target-task inputs while constraining change from pretrained weights: minimize EXTℓ[∥fℓ(XTℓ,Θsftℓ)−fℓ(XTℓ,Θfusionℓ)∥2]\mathbb{E}_{X_T^\ell}\left[\|f_\ell(X_T^\ell,\Theta^\ell_{\mathrm{sft}})-f_\ell(X_T^\ell,\Theta^\ell_{\mathrm{fusion}})\|^2\right], subject to ∥Θfusionℓ−Θpreℓ∥1<η\|\Theta^\ell_{\mathrm{fusion}}-\Theta^\ell_{\mathrm{pre}}\|_1<\eta. Here XTℓX_T^\ell denotes target-task layer inputs, and η\eta bounds the allowed parameter deviation. The paper first expresses original-task preservation as layer-output consistency, but notes that original-task inputs may be unavailable; it therefore uses closeness to pretrained parameters as a practical proxy. In its explicit layer-wise objective, target-output matching is written as ∥ΘsftℓXTℓ−ΘfusionℓXTℓ∥22\|\Theta^\ell_{\mathrm{sft}}X_T^\ell-\Theta^\ell_{\mathrm{fusion}}X_T^\ell\|_2^2.

  3. Knowl 3 — Patch selection fuses parameter salience and loss sensitivity

    model/method

    For each parameter θm\theta_m in layer ℓ\ell, Model Tailor computes a salience score sΔ=∣θsftm−θprem∣s_\Delta=|\theta^m_{\mathrm{sft}}-\theta^m_{\mathrm{pre}}| and an OBS-inspired sensitivity score sε=(θsftm−θprem)2/(2[H−1]mm)s_\varepsilon=(\theta^m_{\mathrm{sft}}-\theta^m_{\mathrm{pre}})^2/(2[H^{-1}]_{mm}). Here θsftm\theta^m_{\mathrm{sft}} and θprem\theta^m_{\mathrm{pre}} are the fine-tuned and pretrained scalar weights, H=∇Θsftℓ2LTH=\nabla^2_{\Theta^\ell_{\mathrm{sft}}}L_T is the target-loss Hessian at the fine-tuned parameters, and [H−1]mm[H^{-1}]_{mm} is the mmth diagonal entry of its inverse. The two scores are min-max normalized within the layer and combined as s=ωs~Δ+(1−ω)s~εs=\omega\tilde{s}_\Delta+(1-\omega)\tilde{s}_\varepsilon, where ω\omega controls the balance. The binary mask retains parameters above the score cutoff for the chosen patch proportion; the reported 10% settings modify 10% of the parameters in the tuned component. In the reported mask-selection analysis, greater weight on salience favored target-task performance, while greater weight on sensitivity favored pretrained-task performance. The sparsity analysis found improved retention on pretrained tasks as sparsity increased, but declining target-task performance; performance largely plateaued beyond 10%, and the authors interpret gains from 0% to 10% as evidence of redundant fine-tuned parameters.

  4. Knowl 4 — The patch decorator compensates retained weights for excluded parameters

    theoretical result

    When a fine-tuned parameter at index mm is reverted to its pretrained value, Model Tailor uses an inverse-Hessian adjustment to compensate the retained parameters. For the target-task loss Hessian HH at the fine-tuned weights, the optimal local second-order adjustment is ΔΘm∗=−θsftm−θprem[H−1]mmH:,m−1\Delta\Theta_m^*=-\frac{\theta^m_{\mathrm{sft}}-\theta^m_{\mathrm{pre}}}{[H^{-1}]_{mm}}H^{-1}_{:,m}, where H:,m−1H^{-1}_{:,m} is column mm of the inverse Hessian. The adjustment to retained coordinate jj is the jjth component of this vector; the patch decorator applies compensation to the mask-selected coordinates. This result is based on the paper’s local second-order loss analysis and is intended to offset target-task loss caused by restoring excluded weights to pretrained values. In the ablation on Flickr30k, decoration raised InstructBLIP’s reported pretrained-task average, target score, overall average, and H-score from 90.75, 93.21, 91.02, and 91.96 before decoration to 92.94, 94.40, 93.10, and 93.67 after it. For LLaVA, the corresponding measures increased from 58.26, 71.58, 59.74, and 64.24 to 60.18, 75.40, 61.87, and 66.94.

  5. Knowl 5 — Multi-task fusion unions task patches and averages their compensation

    model/method

    For mm target tasks, let Θi\Theta_i be the model fine-tuned on task ii, and let MiM_i and CiC_i be that task’s patch mask and compensation tensor. Model Tailor forms an aggregate mask Magg=⋃i=1mMiM_{\mathrm{agg}}=\bigcup_{i=1}^{m}M_i (elementwise union) and average compensation Cagg=1m∑i=1mCiC_{\mathrm{agg}}=\frac{1}{m}\sum_{i=1}^{m}C_i. It then fuses the task models with the pretrained model as Θfusion=Magg⊙(1m∑i=1mΘi+Cagg)+(I−Magg)⊙Θpre\Theta_{\mathrm{fusion}}=M_{\mathrm{agg}}\odot\left(\frac{1}{m}\sum_{i=1}^{m}\Theta_i+C_{\mathrm{agg}}\right)+(I-M_{\mathrm{agg}})\odot\Theta_{\mathrm{pre}}. The resulting model applies the averaged task-model weights and compensation wherever any task mask selects a parameter, and keeps pretrained weights elsewhere.

  6. Knowl 6 — Evaluation uses two MLLMs, unseen target tasks, and a harmonic-mean score

    experimental setup

    The experiments use InstructBLIP (Vicuna-7B) and LLaVA-1.5 (Vicuna-7B). InstructBLIP is fine-tuned on Flickr30k image captioning or GQA visual question answering and evaluated on COCO, NoCaps (in, near, and out splits), OKVQA, AOKVQA, VQAv2, GQA, and Flickr30k. LLaVA-1.5 is fine-tuned on Flickr30k or OKVQA and evaluated on VQAv2, GQA, VizWiz, SQA, TextVQA, POPE, MM-Bench, MM-Bench-CN, and the target task. The reported patch proportion is 10% for the main comparisons. Fine-tuning involves 288M parameters for InstructBLIP and 2.7B for LLaVA-1.5; the corresponding Model Tailor comparisons modify 28.8M and 273M parameters. Results are summarized with average performance and an H-score, defined as PH=2 Avg(Porigin) Avg(Ptarget)Avg(Porigin)+Avg(Ptarget)P_H=\frac{2\,\mathrm{Avg}(P_{\mathrm{origin}})\,\mathrm{Avg}(P_{\mathrm{target}})}{\mathrm{Avg}(P_{\mathrm{origin}})+\mathrm{Avg}(P_{\mathrm{target}})}, where the two averages are performance over pretrained tasks and target tasks, respectively. This harmonic mean balances retention and adaptation rather than allowing the larger original-task set to dominate.

  7. Knowl 7 — Single-task experiments show better retention–adaptation balance with Model Tailor

    empirical result

    Across the InstructBLIP and LLaVA-1.5 single-task comparisons, Model Tailor uses 10% of the fine-tuned parameters and generally improves the reported balance of original-task retention and target-task performance over standard fine-tuning, DARE, and model grafting. On InstructBLIP fine-tuned for Flickr30k, Model Tailor reports average performance 93.10 and H-score 93.67, compared with 86.08 and 92.12 for fine-tuning, 86.8 and 92.43 for DARE, and 89.25 and 92.78 for grafting. On InstructBLIP fine-tuned for GQA, its average and H-score are 92.02 and 71.00; the corresponding values are 75.31 and 67.42 for fine-tuning, 90.72 and 66.46 for DARE, and 89.90 and 66.16 for grafting. On LLaVA-1.5 fine-tuned for Flickr30k, Model Tailor reaches 61.87 average and 66.94 H-score, versus 56.42 and 63.40 for fine-tuning, 60.12 and 36.64 for DARE, and 61.56 and 60.03 for grafting. For LLaVA-1.5 fine-tuned for OKVQA, it reaches 60.95 and 47.71, versus 47.34 and 46.87 for fine-tuning, 58.31 and 1.64 for DARE, and 58.93 and 41.25 for grafting. The dataset metrics have different scales, so the paper reports the cross-task averages and H-scores alongside individual scores. The experiments also illustrate forgetting after ordinary fine-tuning: LLaVA-1.5’s VizWiz score falls from 50.0 zero-shot to 27.24 after Flickr30k tuning, while InstructBLIP’s COCO score falls from 143.1 zero-shot to 114.2 after Flickr30k tuning.

  8. Knowl 8 — Aggregated patches improve the reported multi-task balance

    empirical result

    In multi-task experiments, Model Tailor combines task-specific masks and compensation values rather than using only one task’s patch. For InstructBLIP, the hybrid Flickr30k–GQA result reports average performance 93.72 and H-score 83.12, compared with 93.10 and 83.12 for the Flickr30k-fine-tuned result and 92.03 and 80.79 for the GQA-fine-tuned result. For LLaVA-1.5, the hybrid Flickr30k–OKVQA result reports average 65.71 and H-score 69.52, compared with 62.72 and 64.80 for the Flickr30k-fine-tuned result and 65.18 and 54.01 for the OKVQA-fine-tuned result. The paper’s multi-task visualization compares single-task patches at 10% sparsity with fusion of two 5% task patches, and describes the fusion as bridging performance dips across tasks while also improving pretrained-task performance.

  9. Knowl 9 — Model Tailor complements LoRA on Flickr30k and OKVQA

    empirical result

    The paper applies Model Tailor after LoRA fine-tuning of LLaVA-1.5, integrating LoRA updates into the model before selecting a 5% patch across approximately 6.7B parameters (about 335M selected parameters). For Flickr30k, LoRA alone reports target score 62.5, average 63.39, and H-score 62.99; Model Tailor plus LoRA reports 81.4, 67.70, and 72.89. For OKVQA, the respective figures are 59.25, 48.45, and 52.48 for LoRA alone, and 59.82, 56.75, and 58.04 for Model Tailor plus LoRA. The paper reports improvements across the evaluated task suite and interprets the result as evidence that post-training patch refinement can complement parameter-efficient fine-tuning.

  10. Knowl 10 — SparseGPT makes the post-training Hessian calculations practical

    empirical result

    The paper uses a SparseGPT-derived layer-wise procedure to make the inverse-Hessian computations for Model Tailor feasible at MLLM scale. It reports reducing the stated computational complexity from O(nd2)O(nd^2) to O(d3)O(d^3), where nn is the number of input samples and dd is the hidden-layer size. In the LLaVA timing comparison, standard fine-tuning takes 4×187204\times18720 GPU-seconds and LoRA training takes 4×40204\times4020 GPU-seconds; Model Tailor’s post-training stage takes 1×12601\times1260 GPU-seconds. The compared post-training times are 1×92401\times9240 for DARE and 1×7801\times780 for grafting. Model Tailor is therefore an additional post-training cost after fine-tuning, not a replacement for the initial task-training stage.

Coverage note — No substantial contributed material is omitted. Proof derivations are excluded, and exhaustive per-benchmark score rows are condensed to the reported cross-task comparisons and representative target-task outcomes.

References

  1. 1.Achiam, J., Adler, S., Agarwal, A., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., and Anderson, P. Nocaps: Novel object captioning at scale. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 8948–8957, 2019.
  3. 3.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736, 2022.
  4. 4.Aljundi, R., Babiloni, F., Elhoseiny, M., Rohrbach, M., and Tuytelaars, T. Memory aware synapses: Learning what (not) to forget. In Proceedings of the European conference on computer vision (ECCV), pp. 139–154, 2018.
  5. 5.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  6. 6.Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. Instructblip: Towards general-purpose vision-language models with instruction tuning. arxiv 2023. arXiv preprint arXiv:2305.06500, 2023.
  7. 7.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  8. 8.Dong, X., Luu, A. T., Lin, M., Yan, S., and Zhang, H. How should pre-trained language models be fine-tuned towards adversarial robustness? Advances in Neural Information Processing Systems, 34:4356–4369, 2021.
  9. 9.Frankle, J. and Carbin, M. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In International Conference on Learning Representations, 2019.
  10. 10.Frantar, E. and Alistarh, D. Optimal brain compression: A framework for accurate post-training quantization and pruning. Advances in Neural Information Processing Systems, 35:4475–4488, 2022.
  11. 11.Frantar, E. and Alistarh, D. Sparsegpt: Massive language models can be accurately pruned in one-shot. In International Conference on Machine Learning, pp. 10323–10337. PMLR, 2023.
  12. 12.Goodfellow, I. J., Mirza, M., Xiao, D., Courville, A., and Bengio, Y. An empirical investigation of catastrophic forgetting in gradient-based neural networks. arXiv preprint arXiv:1312.6211, 2013.
  13. 13.Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6904–6913, 2017.
  14. 14.Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P. Vizwiz grand challenge: Answering visual questions from blind people. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3608–3617, 2018.
  15. 15.Hassibi, B., Stork, D. G., and Wolff, G. J. Optimal brain surgeon and general network pruning. In IEEE international conference on neural networks, pp. 293–299. IEEE, 1993.
  16. 16.He, J., Guo, H., Tang, M., and Wang, J. Continual instruction tuning for large multimodal models. arXiv preprint arXiv:2311.16206, 2023.
  17. 17.Hong, Y., Zhen, H., Chen, P., Zheng, S., Du, Y., Chen, Z., and Gan, C. 3d-llm: Injecting the 3d world into large language models. arXiv preprint arXiv:2307.12981, 2023.
  18. 18.Howard, J. and Ruder, S. Universal language model fine-tuning for text classification. arXiv preprint arXiv:1801.06146, 2018.
  19. 19.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  20. 20.Hubara, I., Nahshan, Y., Hanani, Y., Banner, R., and Soudry, D. Accurate post training quantization with small calibration sets. In International Conference on Machine Learning, pp. 4466–4475. PMLR, 2021.
  21. 21.Hudson, D. A. and Manning, C. D. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 6700–6709, 2019.
  22. 22.Korbak, T., Elsahar, H., Kruszewski, G., and Dymetman, M. Controlling conditional language models without catastrophic forgetting. In International Conference on Machine Learning, pp. 11499–11528. PMLR, 2022.
  23. 23.LeCun, Y., Denker, J., and Solla, S. Optimal brain damage. Advances in neural information processing systems, 2, 1989.
  24. 24.Lee, C., Cho, K., and Kang, W. Mixout: Effective regularization to finetune large-scale pretrained language models. arXiv preprint arXiv:1909.11299, 2019.
  25. 25.Li, B., Zhang, Y., Chen, L., Wang, J., Yang, J., and Liu, Z. Otter: A multi-modal model with in-context instruction tuning. arXiv preprint arXiv:2305.03726, 2023a.
  26. 26.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023b.
  27. 27.Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and Wen, J.-R. Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355, 2023c.
  28. 28.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pp. 740–755. Springer, 2014.
  29. 29.Lin, Y., Tan, L., Lin, H., Zheng, Z., Pi, R., Zhang, J., Diao, S., Wang, H., Zhao, H., Yao, Y., et al. Speciality vs generality: An empirical study on catastrophic forgetting in fine-tuning foundation models. arXiv preprint arXiv:2309.06256, 2023.
  30. 30.Liu, H., Li, C., Li, Y., and Lee, Y. J. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744, 2023a.
  31. 31.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023b.
  32. 32.Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., et al. Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281, 2023c.
  33. 33.Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2507–2521, 2022.
  34. 34.Lyu, W., Dong, X., Wong, R., Zheng, S., Abell-Hart, K., Wang, F., and Chen, C. A multimodal transformer: Fusing clinical notes with structured ehr data for interpretable in-hospital mortality prediction. In AMIA Annual Symposium Proceedings, volume 2022, pp. 719. American Medical Informatics Association, 2022a.
  35. 35.Lyu, W., Zheng, S., Ma, T., and Chen, C. A study of the attention abnormality in trojaned berts. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 4727–4741, 2022b.
  36. 36.Lyu, W., Zheng, S., Ling, H., and Chen, C. Backdoor attacks against transformers with attention enhancement. In ICLR 2023 Workshop on Backdoor Attacks and Defenses in Machine Learning, 2023a.
  37. 37.Lyu, W., Zheng, S., Pang, L., Ling, H., and Chen, C. Attention-enhancing backdoor attacks against bert-based models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 10672–10690, 2023b.
  38. 38.Lyu, W., Lin, X., Zheng, S., Pang, L., Ling, H., Jha, S., and Chen, C. Task-agnostic detector for insertion-based backdoor attacks. arXiv preprint arXiv:2403.17155, 2024.
  39. 39.Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, pp. 3195–3204, 2019.
  40. 40.Masana, M., Liu, X., Twardowski, B., Menta, M., Bagdanov, A. D., and Van De Weijer, J. Class-incremental learning: survey and performance evaluation on image classification. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(5):5513–5533, 2022.
  41. 41.Meng, Z., Zhang, J., Yang, C., Zhan, Z., Zhao, P., and WAng, Y. Diffclass: Diffusion-based class incremental learning. arXiv preprint arXiv:2403.05016, 2024.
  42. 42.Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T. Up or down? adaptive rounding for post-training quantization. In International Conference on Machine Learning, pp. 7197–7206. PMLR, 2020.
  43. 43.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  44. 44.Panigrahi, A., Saunshi, N., Zhao, H., and Arora, S. Task-specific skill localization in fine-tuned language models. arXiv preprint arXiv:2302.06600, 2023.
  45. 45.Ritter, H., Botev, A., and Barber, D. Online structured laplace approximations for overcoming catastrophic forgetting. Advances in Neural Information Processing Systems, 31, 2018.
  46. 46.Schwarz, J., Czarnecki, W., Luketina, J., Grabska-Barwinska, A., Teh, Y. W., Pascanu, R., and Hadsell, R. Progress & compress: A scalable framework for continual learning. In International conference on machine learning, pp. 4528–4537. PMLR, 2018.
  47. 47.Schwenk, D., Khandelwal, A., Clark, C., Marino, K., and Mottaghi, R. A-okvqa: A benchmark for visual question answering using world knowledge. In European Conference on Computer Vision, pp. 146–162. Springer, 2022.
  48. 48.Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M. Towards vqa models that can read. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 8317–8326, 2019.
  49. 49.Sung, Y.-L., Yoon, J., and Bansal, M. Ecoflap: Efficient coarse-to-fine layer-wise pruning for vision-language models. arXiv preprint arXiv:2310.02998, 2023.
  50. 50.Tao, X., Hong, X., Chang, X., Dong, S., Wei, X., and Gong, Y. Few-shot class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12183–12192, 2020.
  51. 51.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  52. 52.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023.
  53. 53.Wang, P., Chen, Q., He, X., and Cheng, J. Towards accurate post-training network quantization via bit-split and stitching. In International Conference on Machine Learning, pp. 9847–9856. PMLR, 2020.
  54. 54.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560, 2022a.
  55. 55.Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Selvan Dhanasekaran, A., Naik, A., Stap, D., et al. Benchmarking generalization via in-context instructions on 1,600+ language tasks. arXiv e-prints, pp. arXiv–2204, 2022b.
  56. 56.Wu, T., He, S., Liu, J., Sun, S., Liu, K., Han, Q.-L., and Tang, Y. A brief overview of chatgpt: The history, status quo and potential future development. IEEE/CAA Journal of Automatica Sinica, 10(5):1122–1136, 2023.
  57. 57.Xuhong, L., Grandvalet, Y., and Davoine, F. Explicit inductive bias for transfer learning with convolutional networks. In International Conference on Machine Learning, pp. 2825–2834. PMLR, 2018.
  58. 58.Yang, Y., Yuan, H., Li, X., Lin, Z., Torr, P., and Tao, D. Neural collapse inspired feature-classifier alignment for few-shot class incremental learning. arXiv preprint arXiv:2302.03004, 2023.
  59. 59.Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78, 2014.
  60. 60.Yu, L., Yu, B., Yu, H., Huang, F., and Li, Y. Language models are super mario: Absorbing abilities from homologous models as a free lunch. arXiv preprint arXiv:2311.03099, 2023.
  61. 61.Zhai, Y., Tong, S., Li, X., Cai, M., Qu, Q., Lee, Y. J., and Ma, Y. Investigating the catastrophic forgetting in multimodal large language models. arXiv preprint arXiv:2309.10313, 2023.
  62. 62.Zhang, H., Li, X., and Bing, L. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858, 2023a.
  63. 63.Zhang, J., Chen, C., Zhuang, W., and Lyu, L. Target: Federated class-continual learning via exemplar-free distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4782–4793, 2023b.
  64. 64.Zhang, M., Yuan, J., He, Y., Li, W., Chen, Z., and Kuang, K. MAP: Towards balanced generalization of iid and ood through model-agnostic adapters. In Proceedings of the IEEE/CVF International Conference on Computer Vision, ICCV, pp. 11921–11931, 2023c.
  65. 65.Zhang, M., Li, H., Wu, F., and Kuang, K. Metacoco: A new few-shot classification benchmark with spurious correlation. In International Conference on Learning Representations, ICLR, 2024.
  66. 66.Zhang, P., Wang, X. D. B., Cao, Y., Xu, C., Ouyang, L., Zhao, Z., Ding, S., Zhang, S., Duan, H., Yan, H., et al. Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv preprint arXiv:2309.15112, 2023d.
  67. 67.Zhang, R., Han, J., Zhou, A., Hu, X., Yan, S., Lu, P., Li, H., Gao, P., and Qiao, Y. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199, 2023e.
  68. 68.Zhang, T., Wu, F., Katiyar, A., Weinberger, K. Q., and Artzi, Y. Revisiting few-sample bert fine-tuning. arXiv preprint arXiv:2006.05987, 2020.
  69. 69.Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023a.
  70. 70.Zhu, D., Li, Y., Shao, Y., Hao, J., Wu, F., Kuang, K., Xiao, J., and Wu, C. Generalized universal domain adaptation with generative flow networks. In ACM International Conference on Multimedia (MM) 2023, 2023b.
  71. 71.Zhu, D., Li, Y., Yuan, J., Li, Z., Kuang, K., and Wu, C. Universal domain adaptation via compressive attention matching. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 6974–6985, 2023c.
  72. 72.Zhu, D., Li, Y., Zhang, M., Yuan, J., Liu, J., Kuang, K., and Wu, C. Bridging the gap: neural collapse inspired prompt tuning for generalization under class imbalance. arXiv preprint arXiv:2306.15955, 2023d.
  73. 73.Zhu, L., Chen, T., Ji, D., Ye, J., and Liu, J. Llafs: When large-language models meet few-shot segmentation. arXiv preprint arXiv:2311.16926, 2023e.
  74. 74.Zhu, L., Ji, D., Chen, T., Xu, P., Ye, J., and Liu, J. Ibd: Alleviating hallucinations in large vision-language models via image-biased decoding. arXiv preprint arXiv:2402.18476, 2024.

Citation

MLA
Zhu, D., et al. “Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2402.12048v1.
APA
Zhu, D., Sun, Z., Li, Z., Shen, T., Yan, K., Ding, S., Kuang, K., & Wu, C. (2024). Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models. arXiv. http://arxiv.org/abs/2402.12048v1
Chicago
Zhu, D., Z. Sun, Z. Li, et al. 2024. “Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models”. arXiv. http://arxiv.org/abs/2402.12048v1.
Harvard
Zhu, D. et al. (2024) “Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.12048v1.
Vancouver
1. Zhu D, Sun Z, Li Z, Shen T, Yan K, Ding S, Kuang K, Wu C (2024) Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models. arXiv

BibTeX

@article{zhu2024model,
  title = {Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models},
  author = {Zhu, Didi and Sun, Zhongyi and Li, Zexi and Shen, Tao and Yan, Ke and Ding, Shouhong and Kuang, Kun and Wu, Chao},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.12048v1},
  eprint = {2402.12048}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/