Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models
Didi ZhuZhongyi SunZexi LiTao ShenKe YanShouhong DingChao WuKun Kuang
Proposes Model Tailor, a parameter-efficient post-training method that updates fewer than ten percent of model parameters through sparse masking and Hessian-based compensation, preventing catastrophic forgetting in multi-modal large language models while maintaining performance on both original and target tasks.
Adapting multimodal artificial intelligence models to new downstream applications often introduces a critical operational challenge known as catastrophic forgetting. When multimodal large language models—which integrate visual perception with complex language reasoning—are fine-tuned on new tasks such as image captioning or visual question answering, their ability to perform their original, foundational tasks frequently degrades by a substantial margin. Existing remedies designed for smaller models or text-only systems fail to address the complex cross-modal architecture and computational demands of large multimodal systems. Full retraining across all tasks remains computationally expensive, while standard parameter-efficient methods such as Low-Rank Adaptation continue to retain redundant parameter updates that overwrite pre-existing capabilities.
The article demonstrates and evaluates a post-training framework named Model Tailor, which aims to efficiently adapt multimodal models to target tasks while preserving their baseline proficiency. The core objective is to integrate task-specific modifications into pre-trained models by selectively replacing no more than 10% of the model parameters and mathematically compensating the retained weights to maintain high performance across both old and new domains.
The researchers evaluated this framework on two standard multimodal systems: InstructBLIP, focusing on adapting its 288-million-parameter vision-language connector, and LLaVA-1.5, focusing on adapting both its connector and deep language layers totaling 2.7 billion parameters. The method was tested across a diverse benchmark suite, fine-tuning on datasets such as Flickr30k, GQA, and OKVQA, and assessing performance retention across broad evaluation sets including COCO, VQAv2, VizWiz, and MM-Bench. The approach breaks down the optimization into layer-wise sub-tasks using second-order approximations, selecting high-impact parameters via a hybrid scoring mechanism that balances parameter change and task sensitivity, and adjusting retained weights using an inverse Hessian compensation method.
The findings confirm that Model Tailor preserves original task capabilities while successfully acquiring new skills. Across standard single-task benchmarks, the framework maintained approximately 99% of pre-trained baseline effectiveness while reaching roughly 97% of the performance achieved by standard full fine-tuning on the new target tasks. By contrast, conventional fine-tuning caused severe degradation; for example, LLaVA's accuracy on the VizWiz benchmark dropped from 50.0% to 27.24% following fine-tuning on Flickr30k. When tested in multi-task scenarios, the framework smoothly combined distinct task adjustments without suffering from performance drops across alternating objectives. Furthermore, when combined with Low-Rank Adaptation on LLaVA-1.5, Model Tailor achieved an 18.9-point gain on Flickr30k and improved the overall balanced evaluation score by 9.9%, demonstrating that it effectively removes extraneous parameter updates.
These results show that engineering teams can reliably specialize multimodal models for domain-specific deployments without sacrificing general intelligence or maintaining disconnected models for separate workflows. By modifying only a small fraction of parameters in a rapid post-training step, organizations can reduce the computing costs, infrastructure requirements, and deployment risks associated with maintaining separate large models. The post-training process requires minimal extra computation—taking roughly 21 minutes on a single graphics processing unit during testing—making it an operationally practical addition to standard training pipelines.
Organizations deploying multimodal foundation models should consider adopting sparse parameter selection and compensation strategies like Model Tailor over standard unconstrained fine-tuning. Engineering teams should also combine this post-training adjustment with parameter-efficient techniques such as Low-Rank Adaptation to prune redundant updates. While the framework relies on layer-wise approximations and a recommended sparsity budget near 10%, its strong performance across both single-task and multi-task settings offers high confidence for integrating specialized visual and reasoning capabilities into existing large model deployments.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Learning without Forgetting establishes the core stability–plasticity problem and a foundational approach to preserving old capabilities while fine-tuning for new tasks.
- Paper: Memory Aware Synapses: Learning what (not) to forget, Rahaf Aljundi et al. (2017). Memory Aware Synapses introduces parameter-importance scoring to protect knowledge during updates, a key conceptual precursor to Model Tailor’s selective parameter retention.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). This unified account of parameter-efficient transfer learning clarifies methods such as LoRA that Model Tailor evaluates and combines with its sparse update strategy.
- Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). This long-horizon study extends forgetting mitigation from multimodal post-training to persistent language-model learning, testing how protective mechanisms compose across many sequential updates.
- Paper: Self-Distillation Enables Continual Learning, Idan Shenfeld et al. (2026). Self-Distillation Fine-Tuning continues the effort to preserve prior capabilities during adaptation, applying a teacher–student learning process to continual skill and knowledge acquisition.
