EMR-Merging: Tuning-Free High-Performance Model Merging

Chenyu HuangPeng YeTao ChenTong HeXiangyu YueWanli Ouyang

article2024NeurIPS98 citations

Proposes a tuning-free model merging technique that constructs a unified base model paired with lightweight task-specific masks and scaling factors, matching multi-task training performance across vision, language, and multimodal domains without needing extra data or optimization.

Listen

The widespread adoption of foundation models fine-tuned across diverse downstream tasks has resulted in an exponential growth of task-specific model checkpoints. Deploying and maintaining separate large models for every individual task incurs prohibitive storage, computational, and infrastructure costs. While multi-task training can unify capabilities into a single model, it demands substantial computational resources and direct access to full training datasets, which is often infeasible due to data privacy constraints. Consequently, model merging—combining multiple fine-tuned models directly at the parameter level without additional training data—has emerged as a vital path forward. However, conventional merging methods face a critical dilemma: they either suffer substantial performance degradation when forced into a single set of weights or require resource-intensive hyperparameter tuning and access to validation data.

The article addresses this fundamental limitation by proposing and evaluating ELECT, MASK & RESCALE-MERGING (EMR-MERGING), a tuning-free model merging framework. The primary objective is to demonstrate that combining a unified model representation with lightweight, task-specific parameter modulators allows a system to retain individual task performance without requiring retraining, hyperparameter optimization, or access to task datasets.

To establish its findings, the research evaluates EMR-MERGING across a wide variety of modalities, architectures, and benchmarks. The approach first extracts a unified task vector by electing the dominant sign direction and maximum magnitude across model parameters. It then generates compact, 1-bit binary masks to align direction and scalar rescalers to align magnitude for each specific task. The authors validate this approach on standard benchmarks as well as expanded test beds, including image classification across 8 to 30 Vision Transformer models, natural language understanding tasks using RoBERTa and GPT-2 architectures, parameter-efficient fine-tuning adapter modules, and multi-modal vision-language models like BEiT3.

The evaluation yields several key findings in order of operational importance. First, EMR-MERGING dramatically outperforms existing merging techniques without requiring any tuning data. On standard 8-task computer vision benchmarks, it improves average accuracy by 7.6% over the best competing method on ViT-B/32 and achieves performance comparable to dedicated multi-task training (88.7% versus 88.9%). Second, the framework scales robustly to large task counts: in a 30-task vision benchmark, conventional methods suffered severe degradation, dropping to between 37.5% and 68.1% accuracy, whereas EMR-MERGING maintained 89.5% accuracy, staying within 3.5% of individually fine-tuned models (93.0%). Third, the approach proves broadly applicable across domains, delivering top performance on language benchmarks (outperforming previous GPT-2 merging baselines by over 10%) and vision-language multi-modal tasks. Finally, ablation analyses confirm that the masking and rescaling components can be integrated into other task-vector merging methods to boost their performance by 5.0% to 6.8%.

These findings have direct practical implications for machine learning operations and deployment. Organizations can consolidate dozens of specialized models into a single base model paired with ultra-lightweight modulators, substantially reducing storage footprints and deployment complexity. Because 1-bit masks require 32 times less storage than standard 32-bit model weights and rescalers are single scalar values, the incremental storage overhead per task is minimal. Furthermore, eliminating the need for validation datasets removes compliance and privacy bottlenecks associated with data sharing, accelerating model integration timelines and lowering compute costs.

Organizations seeking to streamline multi-model architectures should consider piloting EMR-MERGING as an alternative to training unified multi-task models or storing multiple separate checkpoints. The primary operational trade-off is the requirement to store and apply task-specific 1-bit modulators at inference time rather than relying entirely on a single static weight set. However, decision-makers should note key technical boundaries: the method assumes that all merged models originate from the same pre-trained base model, meaning it cannot merge models trained from scratch or those utilizing differing underlying architectures. Further research is recommended to explore merging across heterogeneous model designs and integrating the framework with low-bit quantization.

Cover for EMR-Merging: Tuning-Free High-Performance Model Merging

Abstract

The success of pretrain-finetune paradigm brings about the release of numerous model weights. In this case, merging models finetuned on different tasks to enable a single model with multi-task capabilities is gaining increasing attention for its practicability. Existing model merging methods usually suffer from (1) significant performance degradation or (2) requiring tuning by additional data or training. In this paper, we rethink and analyze the existing model merging paradigm. We discover that using a single model’s weights can hardly simulate all the models’ performance. To tackle this issue, we propose ELECT, MASK & RESCALE-MERGING (EMR-MERGING). We first (a) elect a unified model from all the model weights and then (b) generate extremely light-weight task-specific modulators, including masks and rescalers, to align the direction and magnitude between the unified model and each specific model, respectively. EMR-MERGING is tuning-free, thus requiring no data availability or any additional training while showing impressive performance. We find that EMR-MERGING shows outstanding performance compared to existing merging methods under different classical and newly-established settings, including merging different numbers of vision models (up to 30), NLP models, PEFT models, and multi-modal models.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Motivation
  • 3.2 Elect, Mask & Rescale-Merging
  • 3.3 Theoretical analysis
  • 3.4 Empirical analysis
  • 4 Experiment Validation
  • 4.1 Merging vision models
  • 4.1.1 Merging 8 ViTs.
  • 4.1.2 Merging 30 ViTs.
  • 4.2 Merging language models
  • 4.2.1 Merging fully finetuned RoBERTa models
  • 4.2.2 Merging fully finetuned GPT-2 models
  • 4.2.3 Merging PEFT models
  • 4.3 Merging multi-modal models
  • 4.4 Merging different number of models
  • 4.5 Ablation Study
  • 5 Conclusion
  • 6 Acknowledgement
  • References
  • A Algorithm flow of EMR-Merging
  • B Theoretical analyses
  • C Baseline Methods
  • D More experimental results
  • D.1 Merging ViT-B/16 models on 8 tasks
  • D.2 Merging ViT-B/32 models on 9 tasks (ImageNet-1K added)
  • D.3 DARE’s experimental results and causes
  • D.4 Results under different hyper-paramerter settings
  • D.5 Detailed information for merging different number of models
  • D.6 Sparsity of masks and values of rescalers.
  • E More visualization results
  • F Configuration of Fig. and Fig.
  • G Limitations and future works

Knowls

  1. Knowl 1 — EMR-Merging Framework and Algorithm

    algorithm

    ELECT, MASK & RESCALE-MERGING (EMR-MERGING) is a tuning-free model merging technique that constructs a shared unified task vector alongside lightweight, task-specific 1-bit binary masks and scalar rescalers to align parameter directions and magnitudes without requiring validation data, tuning, or retraining.

    Input: Fine-tuned model checkpoints W1,W2,…,WN∈RdW_1, W_2, \dots, W_N \in \mathbb{R}^d for NN tasks, pre-trained base model Wpre∈RdW_{\text{pre}} \in \mathbb{R}^d
    Output: Unified task vector τuni∈Rd\tau_{\text{uni}} \in \mathbb{R}^d, task-specific binary masks M1,…,MN∈{0,1}dM_1, \dots, M_N \in \{0, 1\}^d, task-specific scalar rescalers λ1,…,λN∈R\lambda_1, \dots, \lambda_N \in \mathbb{R}
    for t=1t = 1 to NN do
        τt=Wt−Wpre\tau_t = W_t - W_{\text{pre}}
    end
    γuni=sgn⁡(∑t=1Nτt)\gamma_{\text{uni}} = \operatorname{sgn}\left(\sum_{t=1}^N \tau_t\right)
    ϵuni=0d\epsilon_{\text{uni}} = \mathbf{0}_d
    for t=1t = 1 to NN do
        for p=1p = 1 to dd do
            if γunip⋅τtp>0\gamma_{\text{uni}}^p \cdot \tau_t^p > 0 then
                ϵunip=max⁡(ϵunip,∣τtp∣)\epsilon_{\text{uni}}^p = \max\left(\epsilon_{\text{uni}}^p, |\tau_t^p|\right)
            end
        end
    end
    τuni=γuni⊙ϵuni\tau_{\text{uni}} = \gamma_{\text{uni}} \odot \epsilon_{\text{uni}}
    for t=1t = 1 to NN do
        for p=1p = 1 to dd do
            Mtp=I(τtp⋅τunip>0)M_t^p = \mathbb{I}\left(\tau_t^p \cdot \tau_{\text{uni}}^p > 0\right)
        end
        λt=∑p=1d∣τtp∣∑p=1d∣Mtp⋅τunip∣\lambda_t = \frac{\sum_{p=1}^d |\tau_t^p|}{\sum_{p=1}^d |M_t^p \cdot \tau_{\text{uni}}^p|}
    end

    At inference time for task tt, the task-adapted weights W^t\hat{W}_t are reconstructed via: W^t=Wpre+λt⋅(Mt⊙τuni)\hat{W}_t = W_{\text{pre}} + \lambda_t \cdot (M_t \odot \tau_{\text{uni}}) where ⊙\odot denotes the element-wise Hadamard product.

  2. Knowl 2 — Distance Minimization Guarantee for Masking and Rescaling

    theoretical result

    Let τi=Wi−Wpre∈Rd\tau_i = W_i - W_{\text{pre}} \in \mathbb{R}^d be the task vector of model WiW_i for task i∈{1,…,N}i \in \{1, \dots, N\}, and let τuni∈Rd\tau_{\text{uni}} \in \mathbb{R}^d be the elected unified task vector. The average squared L2L_2 distance between the individual task vectors and τuni\tau_{\text{uni}} is: Dis=1N∑i=1N∥τi−τuni∥2Dis = \frac{1}{N} \sum_{i=1}^N \|\tau_i - \tau_{\text{uni}}\|^2

    Applying the task-specific sign-alignment mask Mi=I(τi⊙τuni>0)M_i = \mathbb{I}(\tau_i \odot \tau_{\text{uni}} > 0) decomposes the distance as: Dis=DisM+1N∑i=1N∥(1−Mi)⊙∣τuni∣∥2Dis = Dis^M + \frac{1}{N} \sum_{i=1}^N \|(1 - M_i) \odot |\tau_{\text{uni}}|\|^2 where DisM=1N∑i=1N∥τi−Mi⊙τuni∥2Dis^M = \frac{1}{N} \sum_{i=1}^N \|\tau_i - M_i \odot \tau_{\text{uni}}\|^2, establishing that DisM≤DisDis^M \le Dis.

    Further applying a task-specific scalar rescaler λi>0\lambda_i > 0 yields the distance: DisM,λ=1N∑i=1N∥τi−λi⋅Mi⊙τuni∥2Dis^{M,\lambda} = \frac{1}{N} \sum_{i=1}^N \|\tau_i - \lambda_i \cdot M_i \odot \tau_{\text{uni}}\|^2 Setting the partial derivative ∂DisM,λ∂λi=0\frac{\partial Dis^{M,\lambda}}{\partial \lambda_i} = 0 yields the optimal analytic rescaler: λi=∑p=1d∣τip∣∑p=1d∣(Mi⊙τuni)p∣\lambda_i = \frac{\sum_{p=1}^d |\tau_i^p|}{\sum_{p=1}^d |(M_i \odot \tau_{\text{uni}})^p|} which guarantees that DisM,λ≤DisM≤DisDis^{M,\lambda} \le Dis^M \le Dis.

  3. Knowl 3 — Multi-Task Vision Model Merging on Eight Benchmarks

    data/table

    When merging fine-tuned CLIP ViT-B/32 and ViT-L/14 models across eight vision benchmarks (SUN397, Cars, RESISC45, EuroSAT, SVHN, GTSRB, MNIST, and DTD), EMR-MERGING achieves multi-task accuracy competitive with both individual fine-tuned models and traditional multi-task learning (MTL), without requiring test/validation data or coefficient tuning.

    Methods SUN397 Cars RESISC45 EuroSAT SVHN GTSRB MNIST DTD Avg Acc
    ViT-B/32 Visual Encoder
    Individual 75.3 77.7 96.1 99.7 97.5 98.7 99.7 79.4 90.5
    Traditional MTL 73.9 74.4 93.9 98.2 95.8 98.9 99.5 77.9 88.9
    Weight Averaging 65.3 63.4 71.4 71.7 64.2 52.8 87.5 50.1 65.8
    Fisher Merging 68.6 69.2 70.7 66.4 72.9 51.1 87.9 59.9 68.3
    RegMean 65.3 63.5 75.6 78.6 78.1 67.4 93.7 52.0 71.8
    Task Arithmetic 63.8 62.1 72.0 77.6 74.4 65.1 94.0 52.2 70.1
    Ties-Merging 64.8 62.9 74.3 78.9 83.1 71.4 97.6 56.2 73.6
    AdaMerging 64.5 68.1 79.2 93.8 87.0 91.9 97.5 59.1 80.1
    AdaMerging++ 66.6 68.3 82.2 94.2 89.6 89.0 98.3 60.6 81.1
    EMR-MERGING (Ours) 75.2 72.8 93.5 99.5 96.9 98.1 99.6 74.4 88.7
    ViT-L/14 Visual Encoder
    Individual 82.3 92.4 97.4 100.0 98.1 99.2 99.7 84.1 94.2
    Traditional MTL 80.8 90.6 96.3 96.3 97.6 99.1 99.6 84.4 93.5
    Weight Averaging 72.1 81.6 82.6 91.9 78.2 70.7 97.1 62.8 79.6
    Fisher Merging 69.2 88.6 87.5 93.5 80.6 74.8 93.3 70.0 82.2
    RegMean 73.3 81.8 86.1 97.0 88.0 84.2 98.5 60.8 83.7
    Task Arithmetic 74.1 82.1 86.7 93.8 87.9 86.8 98.9 65.6 84.5
    Ties-Merging 76.5 85.0 89.3 95.7 90.3 83.3 99.0 68.8 86.0
    AdaMerging 79.0 90.3 90.8 96.2 93.4 98.0 99.0 79.9 90.8
    AdaMerging++ 79.4 90.3 91.6 97.4 93.4 97.5 99.0 79.2 91.0
    EMR-MERGING (Ours) 83.2 90.7 96.8 99.7 97.9 99.1 99.7 82.7 93.7
  4. Knowl 4 — Large-Scale 30-Task Vision Benchmark Merging

    data/table

    On a large-scale setting merging 30 distinct fine-tuned ViT-B/16 checkpoints (pre-trained on ImageNet-21k), existing merging techniques experience significant performance collapse, whereas EMR-MERGING retains near-individual task accuracy.

    Method Average Accuracy (%) on 30 Vision Tasks
    Individual Models 93.02
    Weight Averaging 42.52
    Ties-Merging 37.53
    Task Arithmetic 48.89
    AdaMerging 60.25
    RegMean 68.14
    EMR-MERGING (Ours) 89.54

    The 30 tasks evaluated span image domains including MNIST, CIFAR-10, Vegetables, Food-101, Kvasir-v2, Intel-Images, Cars, EuroSAT, Weather, Cats and Dogs, MangoLeafBD, Beans, CIFAR-100, GTSRB, SVHN, Dogs, Fashion-MNIST, Oxford-IIIT-Pet, Landscape, Flowers, STL-10, CUB-200-2011, EMNIST, DTD, RESISC45, SUN397, KenyanFood13, Animal-10N, Garbage Classification, and Fruits-360. While the strongest baseline (RegMean) incurs a 24.88% degradation relative to individual models, EMR-MERGING restricts the degradation to 3.48%.

  5. Knowl 5 — Full Fine-Tuning Language Model Merging on GLUE

    data/table

    Merging fully fine-tuned RoBERTa-base (8 tasks) and GPT-2 (7 tasks) checkpoints on GLUE benchmark datasets demonstrates that EMR-MERGING scales to language models and substantially outperforms prior merging approaches.

    Methods CoLA SST2 MRPC STSB QQP MNLI QNLI RTE Avg
    RoBERTa-Base (8 GLUE Tasks)
    Individual 0.6018 0.9404 0.8922 0.9063 0.9141 0.8720 0.9271 0.7906 0.8431
    Weight Averaging 0.1396 0.6411 0.6936 0.3184 0.7536 0.4219 0.5870 0.5523 0.5134
    RegMean 0.3667 0.9060 0.7574 0.6268 0.8355 0.7002 0.8235 0.5848 0.7001
    Task Arithmetic 0.1878 0.8589 0.7990 0.7403 0.8378 0.5908 0.6967 0.6209 0.6665
    Ties-Merging 0.2048 0.8440 0.8113 0.5819 0.8570 0.6465 0.7481 0.4296 0.6404
    EMR-MERGING 0.3996 0.9335 0.8627 0.8277 0.8972 0.8545 0.8957 0.7437 0.8018
    GPT-2 (7 GLUE Classification Tasks, Accuracy %)
    Individual 76.8 82.1 80.4 - 88.3 89.6 65.3 91.2 82.0
    Weight Averaging 55.0 55.1 51.0 - 57.6 76.7 44.8 52.5 56.1
    Fisher Merging 54.8 58.0 39.5 - 63.3 81.5 49.1 64.7 58.7
    RegMean 61.7 70.4 65.4 - 69.7 78.8 56.0 79.7 68.8
    Task Arithmetic 68.7 68.6 69.6 - 70.5 81.8 47.3 83.6 70.0
    Ties-Merging 68.4 71.4 68.4 - 69.6 82.4 47.7 81.8 70.0
    EMR-MERGING 72.8 81.1 79.2 - 84.8 88.1 66.5 90.3 80.4

    For RoBERTa, CoLA is measured via Matthews correlation coefficient, STS-B by average Pearson/Spearman correlation, and all other tasks by accuracy. For GPT-2, all tasks are evaluated by classification accuracy.

  6. Knowl 6 — Parameter-Efficient Fine-Tuning Module Merging with (IA)3

    data/table

    When merging Parameter-Efficient Fine-Tuning (PEFT) modules using (IA)3(IA)^3 on a T0-3B base model across eleven NLP tasks, EMR-MERGING outperforms all previous merging approaches.

    Methods Valid. RTE CB Wino. WiC WSC COPA H-SWAG Story ANLI-R1 ANLI-R2 ANLI-R3 Avg
    Individual - 82.7 95.8 75.1 71.7 65.3 85.3 44.4 94.9 70.2 46.5 53.0 71.4
    Traditional MTL - 88.6 95.8 75.5 61.1 80.6 94.1 42.3 97.6 70.5 49.8 47.7 73.1
    Fisher Merging ✓ 83.3 83.3 56.7 54.2 58.3 83.1 42.2 94.1 45.9 41.0 42.2 62.2
    RegMean ✓ 81.2 58.3 53.8 55.2 53.5 80.9 40.1 92.5 43.3 39.2 40.2 58.0
    Task Arithmetic ✓ 74.1 83.3 62.8 49.1 49.3 87.5 41.5 95.3 60.8 49.4 50.0 63.9
    Ties-Merging ✓ 78.0 83.3 67.9 57.6 59.7 81.7 42.8 90.3 66.9 51.3 51.1 66.4
    Weight Averaging ×\times 81.2 58.3 53.8 55.2 53.5 80.9 40.1 92.5 43.3 39.2 40.2 58.0
    Task Arithmetic ×\times 76.5 79.2 57.7 51.6 51.4 66.2 31.4 81.5 59.8 47.5 48.2 59.2
    Ties-Merging ×\times 81.2 87.5 60.8 59.9 58.3 80.2 42.6 91.1 58.1 46.5 47.4 64.9
    EMR-MERGING ×\times 81.8 87.5 66.6 56.1 65.3 82.4 44.7 93.6 65.7 43.8 50.8 67.1

    EMR-MERGING delivers a 67.1% average accuracy without accessing any validation data for hyper-parameter selection, beating validation-free Ties-Merging (64.9%) and Task Arithmetic (59.2%), and also surpassing the best validation-dependent baseline (Ties-Merging at 66.4%).

  7. Knowl 7 — Multi-Modal Model Merging with BEiT3

    data/table

    EMR-MERGING generalizes to multi-modal vision-language architectures, evaluated by merging fine-tuned BEiT3-base checkpoints across five heterogeneous tasks.

    Methods COCO-Retr. COCO-Captioning IN-1k NLVR2 VQAv2
    Acc (%) BLEU4 CIDEr METEOR ROUGE-L Acc (%) Acc (%) Acc (%)
    Individual 0.8456 0.394 1.337 0.311 0.601 0.8537 0.7765 0.8439
    Weight Averaging 0.1893 0.031 0.001 0.115 0.159 0.6771 0.2800 0.6285
    Task Arithmetic 0.3177 0.033 0.000 0.118 0.176 0.7081 0.3809 0.6933
    Ties-Merging 0.3929 0.029 0.001 0.108 0.167 0.6978 0.3206 0.6717
    EMR-MERGING 0.7946 0.289 1.060 0.272 0.534 0.7742 0.7475 0.7211

    The evaluation covers COCO-Retrieval (Image-Text Retrieval), COCO-Captioning (evaluated via BLEU4, CIDEr, METEOR, ROUGE-L), ImageNet-1k Classification, NLVR2 (Visual Reasoning), and VQAv2 (Visual Question Answering). Baseline merging methods suffer severe degradation on generative captioning (CIDEr ≤0.001\le 0.001) and retrieval (accuracy ≤0.3929\le 0.3929), whereas EMR-MERGING retains high multi-modal generative and discriminative capacities.

  8. Knowl 8 — Portability of Mask and Rescale Modulators Across Merging Techniques

    empirical result

    The task-specific Masking and Rescaling (M&R) modulators proposed in EMR-MERGING are modular and can function as a universal post-processing enhancement on top of other task vector merging algorithms.

    When merging eight ViT-B/32 vision models:

    • Task Arithmetic average accuracy improves from 70.1% to 76.9% (+6.8% gain) when augmented with M&R.
    • Ties-Merging average accuracy improves from 73.6% to 78.6% (+5.0% gain) when augmented with M&R.
    • AdaMerging++ average accuracy improves from 81.1% to 87.7% (+6.6% gain) when augmented with M&R.

    However, full EMR-MERGING—which uses the proposed sign-and-magnitude electing procedure as its base—reaches the highest performance at 88.7% average accuracy, outperforming all other combinations.

  9. Knowl 9 — Ablation of Elect, Mask, and Rescale Components

    data/table

    Ablating the three stages of EMR-MERGING (Elect, Mask, Rescale) on the 8-task ViT-B/32 benchmark demonstrates that all three components are essential for preventing performance degradation.

    Configuration SUN397 Cars RESISC45 EuroSAT SVHN GTSRB MNIST DTD Avg Acc
    Elect Only 31.7 34.7 51.8 65.9 85.7 64.0 98.2 42.2 59.3
    Elect + Rescale 58.2 57.2 69.1 81.6 85.2 73.0 98.4 52.2 71.9 (+12.6)
    Elect + Mask 70.7 65.9 92.2 98.7 96.9 97.6 99.6 72.3 86.8 (+27.5)
    Elect + Mask + Rescale 75.2 72.8 93.5 99.5 96.9 98.1 99.6 74.4 88.7 (+29.4)

    Electing alone without modulators achieves 59.3% accuracy due to directional conflicts and magnitude distortion. Direction alignment via 1-bit masks provides the largest single improvement (+27.5%), and combining masking with magnitude rescaling achieves 88.7% (+29.4% over Elect alone).

  10. Knowl 10 — Architectural Limitations and Storage Overhead of EMR-Merging

    limitation

    EMR-MERGING has two primary limitations:

    1. Storage Overhead: Unlike methods that output a single static set of weights, EMR-MERGING requires storing a 1-bit binary mask Mi∈{0,1}dM_i \in \{0, 1\}^d and a single float scalar λi∈R\lambda_i \in \mathbb{R} per task. While a 1-bit mask occupies 1/321/32 (approx. 3.125%) of the storage of a 32-bit model checkpoint—making NN masks substantially smaller than NN individual model checkpoints—it still scales linearly in task count NN.
    2. Pre-Training Dependency: Because EMR-MERGING relies on task vectors τi=Wi−Wpre\tau_i = W_i - W_{\text{pre}}, it strictly requires all models to originate from the identical pre-trained initialization and cannot merge models trained from scratch or models with mismatched architectures.

Coverage note — None was omitted. All key methodological formulas, the core algorithm, theoretical bounds, comprehensive empirical benchmarks (8-task ViT, 30-task ViT, RoBERTa, GPT-2, PEFT, BEiT3), component ablations, and stated limitations are fully represented.

Citation

MLA
Huang, C., et al. “EMR-Merging: Tuning-Free High-Performance Model Merging”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 122741–69, https://proceedings.neurips.cc/paper_files/paper/2024/file/dda5cac5272a9bcd4bc73d90bc725ef1-Paper-Conference.pdf.
APA
Huang, C., Ye, P., Chen, T., He, T., Yue, X., & Ouyang, W. (2024). EMR-Merging: Tuning-Free High-Performance Model Merging. Advances in Neural Information Processing Systems, 37, 122741–122769. https://proceedings.neurips.cc/paper_files/paper/2024/file/dda5cac5272a9bcd4bc73d90bc725ef1-Paper-Conference.pdf
Chicago
Huang, C., P. Ye, T. Chen, T. He, X. Yue, and W. Ouyang. 2024. “EMR-Merging: Tuning-Free High-Performance Model Merging”. Advances in Neural Information Processing Systems 37: 122741–69. https://proceedings.neurips.cc/paper_files/paper/2024/file/dda5cac5272a9bcd4bc73d90bc725ef1-Paper-Conference.pdf.
Harvard
Huang, C. et al. (2024) “EMR-Merging: Tuning-Free High-Performance Model Merging”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 122741–122769. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/dda5cac5272a9bcd4bc73d90bc725ef1-Paper-Conference.pdf.
Vancouver
1. Huang C, Ye P, Chen T, He T, Yue X, Ouyang W (2024) EMR-Merging: Tuning-Free High-Performance Model Merging. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 122741–122769

BibTeX

@inproceedings{huang2024emr,
  title = {EMR-Merging: Tuning-Free High-Performance Model Merging},
  author = {Huang, Chenyu and Ye, Peng and Chen, Tao and He, Tong and Yue, Xiangyu and Ouyang, Wanli},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {122741-122769},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/dda5cac5272a9bcd4bc73d90bc725ef1-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors