EMR-Merging: Tuning-Free High-Performance Model Merging
Chenyu HuangPeng YeTao ChenTong HeXiangyu YueWanli Ouyang
Proposes a tuning-free model merging technique that constructs a unified base model paired with lightweight task-specific masks and scaling factors, matching multi-task training performance across vision, language, and multimodal domains without needing extra data or optimization.
The widespread adoption of foundation models fine-tuned across diverse downstream tasks has resulted in an exponential growth of task-specific model checkpoints. Deploying and maintaining separate large models for every individual task incurs prohibitive storage, computational, and infrastructure costs. While multi-task training can unify capabilities into a single model, it demands substantial computational resources and direct access to full training datasets, which is often infeasible due to data privacy constraints. Consequently, model merging—combining multiple fine-tuned models directly at the parameter level without additional training data—has emerged as a vital path forward. However, conventional merging methods face a critical dilemma: they either suffer substantial performance degradation when forced into a single set of weights or require resource-intensive hyperparameter tuning and access to validation data.
The article addresses this fundamental limitation by proposing and evaluating ELECT, MASK & RESCALE-MERGING (EMR-MERGING), a tuning-free model merging framework. The primary objective is to demonstrate that combining a unified model representation with lightweight, task-specific parameter modulators allows a system to retain individual task performance without requiring retraining, hyperparameter optimization, or access to task datasets.
To establish its findings, the research evaluates EMR-MERGING across a wide variety of modalities, architectures, and benchmarks. The approach first extracts a unified task vector by electing the dominant sign direction and maximum magnitude across model parameters. It then generates compact, 1-bit binary masks to align direction and scalar rescalers to align magnitude for each specific task. The authors validate this approach on standard benchmarks as well as expanded test beds, including image classification across 8 to 30 Vision Transformer models, natural language understanding tasks using RoBERTa and GPT-2 architectures, parameter-efficient fine-tuning adapter modules, and multi-modal vision-language models like BEiT3.
The evaluation yields several key findings in order of operational importance. First, EMR-MERGING dramatically outperforms existing merging techniques without requiring any tuning data. On standard 8-task computer vision benchmarks, it improves average accuracy by 7.6% over the best competing method on ViT-B/32 and achieves performance comparable to dedicated multi-task training (88.7% versus 88.9%). Second, the framework scales robustly to large task counts: in a 30-task vision benchmark, conventional methods suffered severe degradation, dropping to between 37.5% and 68.1% accuracy, whereas EMR-MERGING maintained 89.5% accuracy, staying within 3.5% of individually fine-tuned models (93.0%). Third, the approach proves broadly applicable across domains, delivering top performance on language benchmarks (outperforming previous GPT-2 merging baselines by over 10%) and vision-language multi-modal tasks. Finally, ablation analyses confirm that the masking and rescaling components can be integrated into other task-vector merging methods to boost their performance by 5.0% to 6.8%.
These findings have direct practical implications for machine learning operations and deployment. Organizations can consolidate dozens of specialized models into a single base model paired with ultra-lightweight modulators, substantially reducing storage footprints and deployment complexity. Because 1-bit masks require 32 times less storage than standard 32-bit model weights and rescalers are single scalar values, the incremental storage overhead per task is minimal. Furthermore, eliminating the need for validation datasets removes compliance and privacy bottlenecks associated with data sharing, accelerating model integration timelines and lowering compute costs.
Organizations seeking to streamline multi-model architectures should consider piloting EMR-MERGING as an alternative to training unified multi-task models or storing multiple separate checkpoints. The primary operational trade-off is the requirement to store and apply task-specific 1-bit modulators at inference time rather than relying entirely on a single static weight set. However, decision-makers should note key technical boundaries: the method assumes that all merged models originate from the same pre-trained base model, meaning it cannot merge models trained from scratch or those utilizing differing underlying architectures. Further research is recommended to explore merging across heterogeneous model designs and integrating the framework with low-bit quantization.
- Paper: Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time, Mitchell Wortsman et al. (2022). This paper establishes the foundational concept of model soups and weight averaging across fine-tuned checkpoints that parameter-level model merging frameworks like EMR-Merging directly build upon.
- Paper: Task-Specific Skill Localization in Fine-tuned Language Models, Abhishek Panigrahi et al. (2023). This study demonstrates that task skills in fine-tuned models can be localized into compact parameter masks, providing the theoretical motivation and empirical grounding for mask-based parameter modulation in EMR-Merging.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). This paper provides a unified taxonomy of parameter-efficient fine-tuning architectures, clarifying the adapter and modular representations evaluated in multi-task merging settings.
- Paper: Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language Tasks, Wenhui Wang et al. (2023). This work introduces the BEiT-3 foundation architecture, which serves as a core vision-language benchmark model evaluated in the EMR-Merging framework.
- Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). This paper formalizes the problem of directional conflict and task interference in multi-task parameter spaces, which EMR-Merging addresses via dominant sign election and binary masking.
- Paper: Parameter-Efficient Transfer Learning for NLP, Neil Houlsby et al. (2019). This seminal text introduces parameter-efficient adapter modules for NLP models, establishing the adapter components whose merged transfer is analyzed in the source.
- Paper: Localizing Task Information for Improved Model Merging and Compression, Ke Wang et al. (2024). This paper investigates task interference during model merging and extends mask-based modulation by developing TALL-masks and consensus merging for multi-task model fusion and compression.
- Paper: MetaGPT: Merging Large Language Models Using Model Exclusive Task Arithmetic, Yuyan Zhou et al. (2024). This work extends data-free model merging by deriving closed-form scaling solutions via model-exclusive task arithmetic specifically tailored for large language models.
- Paper: LLM Augmented LLMs: Expanding Capabilities through Composition, Rachit Bansal et al. (2024). This study explores an alternative paradigm to weight merging by composing frozen language models through cross-attention to expand multi-domain capabilities.
- Paper: VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding, Yi Xin et al. (2024). This paper applies multi-task parameter adaptation to dense vision understanding, extending unified representation transfer across visual prediction heads.
- Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). This work analyzes composite continual learning architectures over 100 sequential tasks, providing an experimental frontier for merging low-rank adapters over long task horizons.
