Localizing Task Information for Improved Model Merging and Compression

Ke WangNikolaos DimitriadisGuillermo Ortiz-JiménezFrançois FleuretPascal Frossard

article2024ICML107 citations

Reveals that multi-task merged models still preserve single-task information within non-overlapping parameter subsets, introducing TALL-masks and Consensus Merging to dramatically compress multiple fine-tuned checkpoints and boost merging accuracy by filtering out interfering weights.

Listen

Organizations deploying artificial intelligence increasingly rely on pre-trained foundation models fine-tuned on specialized downstream tasks. However, managing and storing dozens of massive, individual fine-tuned models incurs unsustainable computational and storage costs. Merging multiple specialized models into a single multi-task system via weight interpolation offers a potential solution, but existing approaches suffer from steep performance degradation as the number of tasks grows.

The article demonstrates that multi-task degradation is caused by task interference at evaluation rather than the erasure of task knowledge, and it presents two complementary frameworks: TALL-masks for high-ratio model compression and Consensus Merging for improved multi-task fusion.

To evaluate model behavior, the authors conducted empirical analyses across extensive vision and natural language processing benchmarks, ranging from 8 to 20 tasks using standard architectures such as Vision Transformers and T5-large. Through a controlled experiment with non-overlapping parameters, they isolated the root cause of degradation. They then constructed task-specific binary masks that isolate essential parameters from the combined multi-task weight vector and analyzed parameter agreement across tasks to filter out detrimental weights.

The evaluation revealed several key findings. First, task-relevant information is fully preserved after model merging, and applying task-specific binary masks restores over 99% of original fine-tuned accuracy across benchmarks. Second, the compression framework drastically reduces storage footprint; for instance, storing 20 fine-tuned vision models was reduced from 57.0 Gb to 8.2 Gb while retaining 99.7% of performance. Third, parameter agreement analysis identified both "catastrophic" weights (irrelevant to all tasks) and "selfish" weights (beneficial to only one task), which cause interference in combined models. Fourth, eliminating these disruptive parameters via Consensus Merging consistently improved existing fusion baselines, yielding gains such as 4.9% higher absolute accuracy on a 20-task vision benchmark and 6.3% higher accuracy on an 8-task language benchmark.

These findings indicate that teams maintaining multiple task-specific foundation models can achieve dramatic storage cost savings without sacrificing individual model accuracy. In settings requiring a single generalist model, Consensus Merging mitigates performance loss and improves scaling robustness. Practitioners can adopt Consensus Merging alongside existing pipelines since it operates directly on merged weight vectors without requiring retraining.

Organizations should evaluate their deployment requirements: apply TALL-masks when maximum individual task accuracy and aggressive storage compression are required, or apply Consensus Merging when a unified, single-checkpoint deployment is necessary. However, decision-makers should note that optimal parameter pruning thresholds depend on the underlying merging method and model fine-tuning regime. Additional validation is advised when adapting this framework to new architectures or domain-specific data.

Cover for Localizing Task Information for Improved Model Merging and Compression

Abstract

Model merging and task arithmetic have emerged as promising scalable approaches to merge multiple single-task checkpoints to one multi-task model, but their applicability is reduced by significant performance loss. Previous works have linked these drops to interference in the weight space and erasure of important task-specific features. Instead, in this work we show that the information required to solve each task is still preserved after merging as different tasks mostly use non-overlapping sets of weights. We propose TALL-masks, a method to identify these task supports given a collection of task vectors and show that one can retrieve > 99% of the single task accuracy by applying our masks to the multi-task vector, effectively compressing the individual checkpoints. We study the statistics of intersections among constructed masks and reveal the existence of selfish and catastrophic weights, i.e., parameters that are important exclusively to one task and irrelevant to all tasks but detrimental to multi-task fusion. For this reason, we propose Consensus Merging, an algorithm that eliminates such weights and improves the general performance of existing model merging approaches. Our experiments in vision and NLP benchmarks with up to 20 tasks, show that Consensus Merging consistently improves existing approaches. Furthermore, our proposed compression scheme reduces storage from 57Gb to 8.2Gb while retaining 99.7% of original performance.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Task Interference Causes Performance Degradation
  • 4. TALL-masks : Localizing Task-specific Information in multi-task vector
  • 5. Applications
  • 5.1. Compressing Task Vectors
  • 5.2. Improving Model Merging
  • 6. Experiments
  • 6.1. Model Merging
  • 6.2. Compression
  • 6.3. Individual-task Performance
  • 6.4. Performance with Varying Task Combinations
  • 6.5. Effect of Weight-Pruning Threshold
  • 7. Conclusion
  • Impact Statement
  • Acknowledgments
  • References
  • A. Experimental Details
  • A.1. Fine-tuning
  • A.2. Hyper-parameter tuning
  • A.3. Benchmark task contents
  • A.3.1. COMPUTER VISION
  • A.3.2. NATURAL LANGUAGE PROCESSING
  • B. Derivation of Equation 4
  • C. Storage cost calculation
  • D. Additional results
  • D.1. Performance for removing only catastrophic weights
  • D.2. Effect of weight-pruning threshold for 8 and 14 tasks
  • D.3. Performance with Parameter-efficient fine-tuning methods
  • D.4. Distribution of mask agreements with more tasks
  • D.5. Result for ViT-B/16
  • D.6. Full results on individual tasks
  • D.6.1. FULL RESULTS FOR VISION
  • D.6.2. FULL RESULTS FOR NLP
  • E. Sample subset selection protocol

Knowls

  1. Knowl 1 — Task Localization Masks (TALL-masks) Formulation and Construction

    model/method

    Task Localization Masks (TALL-masks) is a method to identify and isolate task-specific parameter subsets from a merged multi-task vector τMTL∈RP\tau_{\text{MTL}} \in \mathbb{R}^P, where PP is the total parameter count. Given a pre-trained model θ0∈RP\theta_0 \in \mathbb{R}^P, fine-tuned models θt∈RP\theta_t \in \mathbb{R}^P, and task vectors τt=θt−θ0\tau_t = \theta_t - \theta_0 for tasks t∈{1,…,T}t \in \{1, \dots, T\}, TALL-masks constructs a binary mask mt∈{0,1}Pm_t \in \{0, 1\}^P such that the reconstructed model θ^t=θ0+mt∘τMTL\hat{\theta}_t = \theta_0 + m_t \circ \tau_{\text{MTL}} approximates the single-task model θt\theta_t.

    The optimal mask mt∗m_t^* minimizes the ℓ1\ell_1 distance between the reconstructed model θ^t\hat{\theta}_t and θt\theta_t:

    mt∗=arg⁡min⁡mt∈{0,1}P∥θ^t−θt∥1=arg⁡min⁡mt∈{0,1}P∥mt∘τMTL−τt∥1m_t^* = \arg\min_{m_t \in \{0, 1\}^P} \|\hat{\theta}_t - \theta_t\|_1 = \arg\min_{m_t \in \{0, 1\}^P} \|m_t \circ \tau_{\text{MTL}} - \tau_t\|_1

    Because the sum decomposes independently over each parameter index n∈{1,…,P}n \in \{1, \dots, P\}, the condition for activating parameter nn (mt(n)=1m_t^{(n)} = 1) reduces to ∣mt(n)τMTL(n)−τt(n)∣≤∣τt(n)∣|m_t^{(n)} \tau_{\text{MTL}}^{(n)} - \tau_t^{(n)}| \le |\tau_t^{(n)}|, which yields:

    mt=I(∣τt∣≥∣τMTL−τt∣⋅λt)m_t = \mathbb{I}\left(|\tau_t| \ge |\tau_{\text{MTL}} - \tau_t| \cdot \lambda_t\right)

    where I(⋅)\mathbb{I}(\cdot) denotes the element-wise indicator function, ∣⋅∣|\cdot| is the element-wise absolute value, and λt>0\lambda_t > 0 is a task-specific scalar hyperparameter tuned over {0.2,0.3,0.4,0.5,0.6}\{0.2, 0.3, 0.4, 0.5, 0.6\} on a held-out validation set to regulate the mask sparsity.

  2. Knowl 2 — Consensus Merging for Multi-Task Model Fusion

    model/method

    Consensus Merging is a train-free model-merging technique that filters a multi-task vector τMTL∈RP\tau_{\text{MTL}} \in \mathbb{R}^P by pruning parameters that cause multi-task interference, producing a single unified model without requiring task masks at inference time.

    Given the set of TT binary TALL-masks {mt}t=1T⊂{0,1}P\{m_t\}_{t=1}^T \subset \{0, 1\}^P computed for each task, a consensus mask mconsensus∈{0,1}Pm_{\text{consensus}} \in \{0, 1\}^P is defined by a threshold k∈{1,…,T}k \in \{1, \dots, T\}:

    mconsensus=I(∑t=1Tmt≥k)m_{\text{consensus}} = \mathbb{I}\left(\sum_{t=1}^T m_t \ge k\right)

    The consensus multi-task vector τconsensus\tau_{\text{consensus}} and the final merged model θmerged\theta_{\text{merged}} are obtained via:

    τconsensus=mconsensus∘τMTL\tau_{\text{consensus}} = m_{\text{consensus}} \circ \tau_{\text{MTL}} θmerged=θ0+ατconsensus\theta_{\text{merged}} = \theta_0 + \alpha \tau_{\text{consensus}}

    where ∘\circ is the Hadamard product, θ0∈RP\theta_0 \in \mathbb{R}^P is the pre-trained model, and α>0\alpha > 0 is a scaling factor tuned on validation data.

    Setting k=2k = 2 prunes catastrophic weights (selected by 0 masks) and selfish weights (selected by exactly 1 mask), retaining only general weights that benefit at least two tasks. Consensus Merging is modular and can be applied on top of multi-task vectors generated by Task Arithmetic (sum of task vectors) or TIES-Merging (trimmed and sign-resolved vectors).

  3. Knowl 3 — Multi-Task Checkpoint Compression via TALL-masks

    model/method

    TALL-masks enables loss-minimized compression of TT independently fine-tuned task models into a compact representation consisting of:

    1. The pre-trained base model θ0\theta_0 (3232 bits per parameter).
    2. A single merged multi-task vector τMTL\tau_{\text{MTL}} (3232 bits per parameter).
    3. TT task-specific binary masks {mt}t=1T\{m_t\}_{t=1}^T (11 bit per parameter per task).

    For inference on task tt, the specialized task model is reconstructed dynamically as:

    θ^t=θ0+mt∘τMTL\hat{\theta}_t = \theta_0 + m_t \circ \tau_{\text{MTL}}

    For a neural network with P′P' trainable parameters and FF frozen parameters (total parameters P=P′+FP = P' + F), storing TT separate fine-tuned checkpoints requires 32(TP′+F)32(T P' + F) bits. Under the TALL-mask compression scheme, the total storage cost is:

    StorageTALL=(64+T)P′+32F bits\text{Storage}_{\text{TALL}} = (64 + T)P' + 32F \text{ bits}

    This achieves a substantial reduction in disk footprint compared to saving separate checkpoints (e.g., reducing 20 Vision Transformer checkpoints from 57.0 Gb to 8.2 Gb) while retaining >99%> 99\% of single-task fine-tuning accuracy.

  4. Knowl 4 — Mask Agreement Fraction and Weight Categories in Model Merging

    definition

    Let {mt}t=1T\{m_t\}_{t=1}^T be a set of binary masks mt∈{0,1}Pm_t \in \{0, 1\}^P computed via TALL-masks for TT tasks across PP model parameters. The mask agreement percentage α({mt}t=1T,n)\alpha(\{m_t\}_{t=1}^T, n) is the fraction of total parameters deemed important by exactly nn tasks:

    α({mt}t=1T,n)=1P∑p=1PI(∑t=1Tmt(p)=n)\alpha(\{m_t\}_{t=1}^T, n) = \frac{1}{P} \sum_{p=1}^P \mathbb{I}\left(\sum_{t=1}^T m_t^{(p)} = n\right)

    Based on agreement count n∈{0,1,…,T}n \in \{0, 1, \dots, T\}, parameters are categorized into four functional classes:

    • Catastrophic weights (n=0n = 0): Parameters deemed unimportant by all task masks; their presence in the multi-task vector introduces interference without benefiting any task.
    • Selfish weights (n=1n = 1): Parameters important to exclusively one task; their presence in a unified multi-task vector degrades performance for all other T−1T-1 tasks.
    • General weights (n≥2n \ge 2): Parameters beneficial to two or more tasks; their utility increases with higher task agreement.
    • Universal weights (n=Tn = T): Parameters essential to all TT tasks simultaneously.
  5. Knowl 5 — Performance Degradation in Merged Models is Caused by Task Interference Rather than Information Erasure

    empirical result

    To evaluate whether accuracy drops in Task Arithmetic stem from information erasure (destruction of task-specific weights during addition) or task interference (competing parameter updates distorting task-specific mappings), a controlled experiment was conducted on an 8-task image classification benchmark (Cars, DTD, EuroSAT, GTSRB, MNIST, RESISC45, SUN397, SVHN) using CLIP ViT-B/32.

    In the controlled setting, trainable weights were partitioned into 8 disjoint subsets of equal size, and each task was fine-tuned exclusively on its allocated subset. By construction, parameter interference and parameter competition were identically zero, and all task-specific weights were perfectly preserved in the multi-task vector τMTL\tau_{\text{MTL}}.

    Fine-tuning method Evaluation Abs. acc. (%) Norm. acc. (%)
    Standard Single-task Fine-tuned 92.8 100.0
    Controlled (Disjoint) Single-task Fine-tuned 89.8 100.0
    Standard Task Arithmetic Merged 71.5 77.0
    Controlled (Disjoint) Task Arithmetic Merged 68.8 76.6

    Task Arithmetic degraded at almost the exact same rate under the controlled non-overlapping condition (retaining 76.6%76.6\% normalized accuracy) as it did under standard fine-tuning (77.0%77.0\% normalized accuracy). This demonstrates that performance degradation in merged models is caused by task interference at evaluation time (updates from other tasks altering discriminant features) rather than erasure of task-specific knowledge.

  6. Knowl 6 — Model Merging and Checkpoint Compression Performance on Vision Benchmarks (ViT-B/32 and ViT-L/14)

    data/table

    The table below presents the evaluation of train-free model merging and checkpoint compression methods across 8-, 14-, and 20-task vision benchmarks using CLIP ViT-B/32 and ViT-L/14 visual encoders. Metrics reported are absolute classification accuracy (%), normalized accuracy relative to fine-tuned single-task models in parentheses (%), and total storage size in gigabits (Gb).

    ViT-B/32 ViT-L/14
    Type Method 8 tasks 14 tasks 20 tasks 8 tasks 14 tasks 20 tasks
    Merging Zero-shot 48.4 (3.6) 57.3 (3.6) 56.1 (3.6) 64.4 (11.0) 68.0 (11.0) 65.1 (11.0)
    Weight averaging 66.5 (72.3) 64.4 (71.2) 61.1 (67.5) 79.4 (83.0) 76.6 (81.0) 71.5 (75.5)
    Task arithmetic 70.8 (76.5) 65.4 (72.2) 60.6 (66.8) 84.8 (88.5) 79.3 (83.8) 74.0 (78.0)
    TIES 75.1 (81.0) 68.0 (74.8) 63.4 (69.9) 86.9 (90.7) 79.5 (84.1) 75.7 (79.8)
    Consensus TA [ours] 75.0 (80.8) 70.4 (77.4) 65.4 (72.0) 86.2 (89.9) 82.2 (86.9) 78.9 (83.2)
    Consensus TIES [ours] 74.8 (80.6) 67.7 (74.5) 63.2 (69.6) 86.9 (90.7) 81.5 (86.1) 76.8 (80.9)
    Compression Fine-tuned (Uncompressed) 92.8 (23.3) 90.9 (40.2) 91.3 (57.0) 95.8 (79.1) 94.3 (137.4) 94.7 (195.8)
    Magnitude Pruning 91.3 (98.4) 85.3 (93.7) 83.4 (91.2) 95.4 (99.6) 91.8 (97.2) 91.2 (96.1)
    Magnitude Masking 86.8 (93.3) 80.7 (88.4) 75.3 (82.1) 94.6 (98.7) 91.6 (97.0) 91.6 (96.5)
    TALL Mask + TA [ours] 92.6 (99.7) 90.1 (99.1) 90.6 (99.2) 95.7 (99.9) 93.1 (98.8) 93.7 (98.9)
    TALL Mask + TIES [ours] 93.0 (100.3) 90.9 (100.0) 91.1 (99.7) 95.9 (100.1) 93.4 (99.0) 93.9 (99.1)

    Consensus TA improves standard Task Arithmetic by up to 4.9%4.9\% absolute accuracy on 20 tasks for both ViT-B/32 and ViT-L/14. In compression, TALL Mask with Task Arithmetic or TIES maintains >99%> 99\% normalized performance across all setups while reducing storage from 195.8 Gb to 26.8 Gb on 20 ViT-L/14 tasks.

  7. Knowl 7 — Model Merging and Checkpoint Compression Performance on NLP Benchmarks with T5-Large

    data/table

    The table below details the performance and storage footprint of model merging and checkpoint compression methods evaluated on T5-large-LM-Adapt across three NLP benchmark setups: 7 NLP classification tasks, 8 Question-Answering (QA) tasks, and all 11 combined tasks. Accuracies are reported as absolute percentage with normalized percentage in parentheses, alongside total storage size in Gb.

    7 NLP tasks 8 QA tasks All 11 tasks
    Type Method Acc. (%) Bits (Gb) Acc. (%) Bits (Gb) Acc. (%) Bits (Gb)
    Merging Zero-shot 44.9 25.1 33.1 25.1 36.9 25.1
    Weight averaging 60.5 (72.7) 25.1 56.4 (69.9) 25.1 55.2 (71.6) 25.1
    Task arithmetic 71.9 (85.3) 25.1 63.8 (79.6) 25.1 63.6 (81.8) 25.1
    TIES 69.6 (83.5) 25.1 62.8 (78.9) 25.1 64.0 (82.6) 25.1
    Consensus TA [ours] 73.5 (87.7) 25.1 68.6 (85.4) 25.1 67.5 (86.8) 25.1
    Consensus TIES [ours] 71.0 (84.2) 25.1 69.1 (85.8) 25.1 66.8 (85.7) 25.1
    Compression Fine-tuned 85.9 169.1 80.7 193.1 78.7 265.1
    Magnitude Pruning 81.6 (93.5) >54.3 70.8 (85.7) >55.1 65.8 (81.4) >57.3
    Magnitude Masking 78.9 (90.7) 54.3 72.5 (88.8) 55.1 69.8 (87.2) 57.3
    TALL Mask + TA [ours] 86.8 (102.2) 54.3 79.6 (98.7) 55.1 76.5 (96.2) 57.3
    TALL Mask + TIES [ours] 83.4 (95.4) 54.3 79.7 (98.8) 55.1 77.4 (97.5) 57.3

    Consensus TA improves upon Task Arithmetic by 4.8%4.8\% on the 8 QA benchmark and by 3.9%3.9\% on the 11-task benchmark. For compression, TALL Mask with Task Arithmetic recovers up to 102.2%102.2\% of individual fine-tuned performance on 7 tasks while compressing model storage from 169.1 Gb to 54.3 Gb.

  8. Knowl 8 — Vision Model Merging and Compression Results on CLIP ViT-B/16

    data/table

    The table below provides classification performance (absolute accuracy and normalized accuracy in parentheses) and storage footprints (Gb) across 8, 14, and 20 image classification tasks for the CLIP ViT-B/16 visual backbone.

    8 tasks 14 tasks 20 tasks
    Type Method Acc. (%) Bits (Gb) Acc. (%) Bits (Gb) Acc. (%) Bits (Gb)
    Merging Zero-shot 55.2 3.6 61.2 3.6 59.7 3.6
    Weight averaging 72.2 (76.6) 3.6 69.4 (74.7) 3.6 65.3 (70.3) 3.6
    Task arithmetic 75.8 (80.2) 3.6 70.5 (75.8) 3.6 65.7 (70.7) 3.6
    TIES 79.7 (84.3) 3.6 73.2 (78.7) 3.6 68.2 (73.3) 3.6
    Consensus TA [ours] 79.2 (83.6) 3.6 74.3 (79.8) 3.6 69.7 (74.9) 3.6
    Consensus TIES [ours] 79.4 (83.9) 3.6 74.1 (79.5) 3.6 68.7 (73.9) 3.6
    Compression Fine-tuned 94.6 22.9 92.8 39.4 93.2 56.0
    Mag. Pruning 93.3 (98.5) >7.0 88.1 (94.7) >7.5 86.5 (92.7) >8.1
    Mag. Masking 89.7 (94.6) 7.0 84.8 (91.1) 7.5 81.6 (87.3) 8.1
    TALL Mask + TA [ours] 94.2 (99.6) 7.0 92.0 (99.2) 7.5 92.5 (99.3) 8.1
    TALL Mask + TIES [ours] 94.6 (99.9) 7.0 92.6 (99.8) 7.5 93.0 (99.8) 8.1

    Consensus TA outperforms Task Arithmetic by 3.4%3.4\% (8 tasks), 3.8%3.8\% (14 tasks), and 4.0%4.0\% (20 tasks) absolute accuracy. TALL Mask compression achieves ≥99.2%\ge 99.2\% normalized accuracy across all task counts while using 8.18.1 Gb versus 56.056.0 Gb for 20 tasks.

  9. Knowl 9 — Consensus Merging with Parameter-Efficient Fine-Tuning ((IA)³)

    data/table

    The table below shows the performance of Consensus Merging when applied to T5-large models fine-tuned with the parameter-efficient fine-tuning (PEFT) method (IA)³ across 7 NLP tasks, 8 QA tasks, and 11 combined tasks. Accuracies are shown as absolute percentage (normalized percentage in parentheses).

    Method 7 NLP tasks 8 QA tasks All 11 tasks
    Zero-shot 44.9 33.1 36.9
    Task arithmetic 67.1 (79.5) 57.6 (74.3) 59.7 (77.7)
    Consensus TA (k=1k = 1) 70.9 (83.6) 60.2 (77.1) 60.8 (79.1)
    Consensus TA (k=2k = 2) 66.6 (78.6) 54.1 (68.1) 55.7 (71.1)
    TIES 66.0 (77.6) 56.8 (72.6) 56.1 (72.5)
    Consensus TIES (k=1k = 1) 71.3 (84.2) 58.7 (74.9) 59.2 (76.5)
    Consensus TIES (k=2k = 2) 67.6 (79.8) 56.3 (71.6) 53.1 (68.8)

    For (IA)³ PEFT checkpoints, setting k=1k = 1 (removing only catastrophic weights) produces consistent improvements over Task Arithmetic (+3.8%+3.8\% on 7 NLP tasks, +2.6%+2.6\% on 8 QA tasks) and TIES (+5.3%+5.3\% on 7 NLP tasks, +1.9%+1.9\% on 8 QA tasks). Setting k=2k = 2 degrades PEFT performance due to the different parameter distribution profile in low-rank/scaling vectors compared to full fine-tuning.

  10. Knowl 10 — Threshold Sensitivity Discrepancy Between Task Arithmetic and TIES in Consensus Merging

    empirical result

    In Consensus Merging, the choice of the weight-pruning threshold kk (the minimum number of task masks in which a parameter must appear to avoid being pruned) exhibits differing optima depending on the underlying merging method:

    • Consensus Task Arithmetic: Achieves optimal multi-task accuracy at k=2k = 2. Task Arithmetic forms simple additive vectors containing a high proportion of selfish weights (n=1n = 1). Eliminating both catastrophic (n=0n = 0) and selfish (n=1n = 1) weights removes the main source of multi-task interference.
    • Consensus TIES: Achieves optimal accuracy at k=1k = 1. The internal sparsification and sign-agreement resolution steps in TIES already shift the parameter distribution by concentrating saliency into fewer parameters and amplifying universal/selfish weights. Pruning selfish weights (k=2k = 2) removes excessively many salient parameters under TIES, making k=1k = 1 (pruning only catastrophic n=0n = 0 parameters) the superior configuration.

Coverage note — Omitted specific subset permutation details (Appendix E) and individual per-task radar plots (Figures 4, 9, 10, 11) because their insights are fully captured by the aggregate empirical tables and the controlled interference experiment.

References

  1. 1.Bayazit, D., Foroutan, N., Chen, Z., Weiss, G., and Bosselut, A. Discovering knowledge-critical subnetworks in pretrained language models. arXiv, 2023. URL http://arxiv.org/abs/2310.03084v1.
  2. 2.Bossard, L., Guillaumin, M., and Van Gool, L. Food-101 – Mining Discriminative Components with Random Forests. In IEEE European Conference on Computer Vision (ECCV), 2014.
  3. 3.Caruana, R. Multitask Learning. Machine Learning, 28(1): 41–75, 1997.
  4. 4.Chaudhari, P. A. A Picture of the Energy Landscape of Deep Neural Networks. PhD thesis, 2018.
  5. 5.Chen, Z., Deng, Y., Wu, Y., Gu, Q., and Li, Y. Towards understanding the mixture-of-experts layer in deep learning. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  6. 6.Cheng, G., Han, J., and Lu, X. Remote sensing image scene classification: Benchmark and state of the art. Proceedings of the IEEE, 2017. URL http://arxiv.org/abs/1703.00121v1.
  7. 7.Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A. Describing textures in the wild. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014. URL http://arxiv.org/abs/1311.3618v2.
  8. 8.Clanuwat, T., Bober-Irizar, M., Kitamoto, A., Lamb, A., Yamamoto, K., and Ha, D. Deep learning for classical japanese literature. arXiv, 2018. URL http://arxiv.org/abs/1812.01718v1.
  9. 9.Coates, A., Ng, A., and Lee, H. An analysis of single-layer networks in unsupervised feature learning. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2011. https://proceedings.mlr.press/v15/coates11a.html.
  10. 10.Cohen, G., Afshar, S., Tapson, J., and Van Schaik, A. EMNIST: Extending MNIST to handwritten letters. In International Joint Conference on Neural Networks (IJCNN), 2017.
  11. 11.Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F. Knowledge Neurons in Pretrained Transformers. In Association for Computational Linguistics (ACL), 2022. URL https://aclanthology.org/2022.acl-long.581.
  12. 12.Davari, M. and Belilovsky, E. Model Breadcrumbs: Scaling Multi-Task Model Merging with Sparse Masks. arXiv, 2023. URL http://arxiv.org/abs/2312.06795v1.
  13. 13.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale. In Advances in Neural Information Processing Systems (NeurIPS), 2022. URL https://openreview.net/forum?id=dXiGWqBoxaD.
  14. 14.Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. Qlora: Efficient finetuning of quantized llms. arXiv, 2023. URL http://arxiv.org/abs/2305.14314v1.
  15. 15.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In North American Chapter of the Association for Computational Linguistics (NAACL), 2019. https://aclanthology.org/N19-1423.
  16. 16.Dimitriadis, N., Frossard, P., and Fleuret, F. Pareto Manifold Learning: Tackling multiple tasks via ensembles of single-task models. In International Conference on Machine Learning (ICML), 2023.
  17. 17.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations (ICLR), 2021. URL http://arxiv.org/abs/2010.11929v2.
  18. 18.Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. A. Essentially No Barriers in Neural Network Energy Landscape. In International Conference on Machine Learning (ICML), 2018. URL http://arxiv.org/abs/1803.00885v5.
  19. 19.Fedus, W., Zoph, B., and Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research (JMLR), 23(120):1–39, 2022. URL http://arxiv.org/abs/2101.03961v3.
  20. 20.Fifty, C., Amid, E., Zhao, Z., Yu, T., Anil, R., and Finn, C. Efficiently Identifying Task Groupings for Multi-Task Learning. arXiv, 2021. URL http://arxiv.org/abs/2109.04617v2.
  21. 21.Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B. Sharpness-Aware Minimization for Efficiently Improving Generalization. In International Conference on Learning Representations (ICLR), 2021. URL http://arxiv.org/abs/2010.01412v3.
  22. 22.Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M. Linear mode connectivity and the lottery ticket hypothesis. In International Conference on Machine Learning (ICML), 2020. URL http://arxiv.org/abs/1912.05671v4.
  23. 23.Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G. Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs. In Advances in Neural Information Processing Systems (NeurIPS), 2018. URL http://arxiv.org/abs/1802.10026v4.
  24. 24.Goodfellow, I. J., Erhan, D., Carrier, P. L., Courville, A., Mirza, M., Hamner, B., Cukierski, W., Tang, Y., Thaler, D., Lee, D.-H., et al. Challenges in representation learning: A report on three machine learning contests. In International Conference on Neural Information Processing (ICONIP), 2013. URL http://arxiv.org/abs/1307.0414v1.
  25. 25.Helber, P., Bischke, B., Dengel, A., and Borth, D. EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification. IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens., 12(7):2217–2226, 2019. URL https://doi.org/10.1109/JSTARS.2019.2918242.
  26. 26.Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning (ICML), 2019. URL http://arxiv.org/abs/1902.00751v2.
  27. 27.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations (ICLR), 2022. https://openreview.net/forum?id=nZeVKeeFYf9.
  28. 28.Huang, L., Bras, R. L., Bhagavatula, C., and Choi, Y. Cosmos QA: Machine reading comprehension with contextual commonsense reasoning. arXiv, 2019. URL http://arxiv.org/abs/1909.00277v2.
  29. 29.Ilharco, G., Wortsman, M., Wightman, R., Gordon, C., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., and Schmidt, L. OpenCLIP, 2021. https://doi.org/10.5281/zenodo.5143773.
  30. 30.Ilharco, G., Wortsman, M., Gadre, S. Y., Song, S., Hajishirzi, H., Kornblith, S., Farhadi, A., and Schmidt, L. Patching open-vocabulary models by interpolating weights. In Advances in Neural Information Processing Systems (NeurIPS), 2022. URL http://arxiv.org/abs/2208.05592v2.
  31. 31.Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A. Editing models with task arithmetic. In International Conference on Learning Representations (ICLR), 2023. https://arxiv.org/abs/2110.08207.
  32. 32.Jiang, W., Lin, B., Shi, H., Zhang, Y., Li, Z., and Kwok, J. T. BYOM: Building Your Own Multi-Task Model For Free, 2024. URL http://arxiv.org/abs/2310.01886v3.
  33. 33.Jin, X., Ren, X., Preotiuc-Pietro, D., and Cheng, P. Dataless Knowledge Fusion by Merging Weights of Language Models. In International Conference on Learning Representations (ICLR), 2023. URL https://openreview.net/forum?id=FCnohuR6AnM.
  34. 34.Khot, T., Clark, P., Guerquin, M., Jansen, P., and Sabharwal, A. QASC: A Dataset for Question Answering via Sentence Composition, 2020. URL http://arxiv.org/abs/1910.11473v2.
  35. 35.Krause, J., Stark, M., Deng, J., and Fei-Fei, L. 3d object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops, 2013.
  36. 36.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images, 2009. https://www.cs.toronto.edu/~kriz/learning-features-2009-TR.pdf.
  37. 37.LeCun, Y. The MNIST database of handwritten digits, 1998. http://yann.lecun.com/exdb/mnist/.
  38. 38.Lester, B., Al-Rfou, R., and Constant, N. The power of scale for parameter-efficient prompt tuning. arXiv, 2021. URL http://arxiv.org/abs/2104.08691v2.
  39. 39.Levesque, H., Davis, E., and Morgenstern, L. The winograd schema challenge. In Thirteenth international conference on the principles of knowledge representation and reasoning, 2012.
  40. 40.Liang, T., Glossner, J., Wang, L., Shi, S., and Zhang, X. Pruning and quantization for deep neural network acceleration: A survey. Neurocomputing, 461:370–403, 2021. URL http://arxiv.org/abs/2101.09671v3.
  41. 41.Lin, K., Tafjord, O., Clark, P., and Gardner, M. Reasoning over paragraph effects in situations. arXiv, 2019. URL http://arxiv.org/abs/1908.05852v2.
  42. 42.Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C. A. Few-shot parameter-efficient finetuning is better and cheaper than in-context learning. In Advances in Neural Information Processing Systems (NeurIPS), 2022. URL http://arxiv.org/abs/2205.05638v2.
  43. 43.Luo, S., Tan, Y., Patil, S., Gu, D., von Platen, P., Passos, A., Huang, L., Li, J., and Zhao, H. Lcm-lora: A universal stable-diffusion acceleration module. arXiv, 2023.
  44. 44.Matena, M. S. and Raffel, C. A. Merging models with fisher-weighted averaging. In Advances in Neural Information Processing Systems (NeurIPS), 2022. URL http://arxiv.org/abs/2111.09832v2.
  45. 45.Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. Reading Digits in Natural Images with Unsupervised Feature Learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011, 2011. URL http://ufldl.stanford.edu/housenumbers/nips2011_housenumbers.pdf.
  46. 46.Nilsback, M.-E. and Zisserman, A. Automated flower classification over a large number of classes. In 2008 Sixth Indian conference on computer vision, graphics & image processing, 2008.
  47. 47.Ortiz-Jimenez, G., Favero, A., and Frossard, P. Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained Models. In Advances in Neural Information Processing Systems (NeurIPS), 2023. URL http://arxiv.org/abs/2305.12827v3.
  48. 48.Panigrahi, A., Saunshi, N., Zhao, H., and Arora, S. Task-specific skill localization in fine-tuned language models. In International Conference on Machine Learning (ICML), 2023. URL http://arxiv.org/abs/2302.06600v2.
  49. 49.Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. Cats and dogs. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012.
  50. 50.Pruksachatkun, Y., Phang, J., Liu, H., Htut, P. M., Zhang, X., Pang, R. Y., Vania, C., Kann, K., and Bowman, S. R. Intermediate-Task Transfer Learning with Pre-trained Language Models: When and Why Does It Work? In Association for Computational Linguistics (ACL), 2020. URL https://aclanthology.org/2020.acl-main.467.
  51. 51.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language Models are Unsupervised Multitask Learners, 2019. https://openai.com/blog/better-language-models/.
  52. 52.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning Transferable Visual Models From Natural Language Supervision. In International Conference on Machine Learning (ICML), 2021. URL http://arxiv.org/abs/2103.00020v1.
  53. 53.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research (JMLR), 21(140):1–67, 2020.
  54. 54.Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., Eccles, T., Bruce, J., Razavi, A., Edwards, A., Heess, N., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., and de Freitas, N. A Generalist Agent. Transactions on Machine Learning Research, 2022. URL https://openreview.net/forum?id=1ikK0kHjvj.
  55. 55.Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Pinto, A. S., Keysers, D., and Houlsby, N. Scaling Vision with Sparse Mixture of Experts. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  56. 56.Rogers, A., Kovaleva, O., Downey, M., and Rumshisky, A. Getting closer to AI complete question answering: A set of prerequisite real tasks. In AAAI Conference on Artificial Intelligence (AAAI), 2020.
  57. 57.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. WinoGrande: An Adversarial Winograd Schema Challenge at Scale. Commun. ACM, 64(9):99–106, 2021. URL http://arxiv.org/abs/1907.10641v2.
  58. 58.Sap, M., Rashkin, H., Chen, D., LeBras, R., and Choi, Y. Socialiqa: Commonsense reasoning about social interactions. arXiv, 2019. URL http://arxiv.org/abs/1904.09728v3.
  59. 59.Sharma, R., Allen, J., Bakhshandeh, O., and Mostafazadeh, N. Tackling the Story Ending Biases in The Story Cloze Test. In Association for Computational Linguistics (ACL), 2018. URL https://aclanthology.org/P18-2119.
  60. 60.Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In International Conference on Learning Representations (ICLR), 2017.
  61. 61.Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In Empirical Methods in Natural Language Processing (EMNLP), 2013. https://aclanthology.org/D13-1170/.
  62. 62.Stallkamp, J., Schlipsing, M., Salmen, J., and Igel, C. The German traffic sign recognition benchmark: a multi-class classification competition. In International Joint Conference on Neural Networks (IJCNN), 2011. https://ieeexplore.ieee.org/document/6033395.
  63. 63.Standley, T., Zamir, A. R., Chen, D., Guibas, L. J., Malik, J., and Savarese, S. Which Tasks Should Be Learned Together in Multi-task Learning? In International Conference on Machine Learning (ICML), 2020. URL http://arxiv.org/abs/1905.07553v4.
  64. 64.Tafjord, O., Gardner, M., Lin, K., and Clark, P. QuaRTz: An Open-Domain Dataset of Qualitative Relationship Questions. In Empirical Methods in Natural Language Processing (EMNLP), 2019. URL https://aclanthology.org/D19-1608.
  65. 65.Tam, D., Bansal, M., and Raffel, C. Merging by Matching Models in Task Parameter Subspaces. Transactions on Machine Learning Research, 2024. URL https://openreview.net/forum?id=qNGo6ghWFB.
  66. 66.Veeling, B. S., Linmans, J., Winkens, J., Cohen, T., and Welling, M. Rotation equivariant CNNs for digital pathology. In International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI), 2018. URL http://arxiv.org/abs/1806.03962v1.
  67. 67.Wortsman, M., Horton, M. C., Guestrin, C., Farhadi, A., and Rastegari, M. Learning Neural Network Subspaces. In International Conference on Machine Learning (ICML), 2021. URL http://arxiv.org/abs/2102.10472v3.
  68. 68.Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In International Conference on Machine Learning (ICML), 2022a. URL http://arxiv.org/abs/2203.05482v3.
  69. 69.Wortsman, M., Ilharco, G., Kim, J. W., Li, M., Kornblith, S., Roelofs, R., Lopes, R. G., Hajishirzi, H., Farhadi, A., Namkoong, H., and Schmidt, L. Robust Fine-Tuning of Zero-Shot Models. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022b. URL http://arxiv.org/abs/2109.01903v3.
  70. 70.Xiao, H., Rasul, K., and Vollgraf, R. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv, 2017. URL http://arxiv.org/abs/1708.07747v2.
  71. 71.Xiao, J., Ehinger, K. A., Hays, J., Torralba, A., and Oliva, A. Sun database: Exploring a large collection of scene categories. International Journal of Computer Vision, 119:3–22, 2016.
  72. 72.Yadav, P., Choshen, L., Raffel, C., and Bansal, M. ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization. arXiv, 2023a. URL http://arxiv.org/abs/2311.13171v1.
  73. 73.Yadav, P., Tam, D., Choshen, L., Raffel, C., and Bansal, M. TIES-Merging: Resolving Interference When Merging Models. In Advances in Neural Information Processing Systems (NeurIPS), 2023b. URL http://arxiv.org/abs/2306.01708v2.
  74. 74.Yang, E., Wang, Z., Shen, L., Liu, S., Guo, G., Wang, X., and Tao, D. AdaMerging: Adaptive Model Merging for Multi-Task Learning. In International Conference on Learning Representations (ICLR), 2024. URL https://openreview.net/forum?id=nZP6NgD3QY.
  75. 75.Yang, Y., Yih, W.-t., and Meek, C. WikiQA: A Challenge Dataset for Open-Domain Question Answering. In Empirical Methods in Natural Language Processing (EMNLP), 2015. URL https://aclanthology.org/D15-1237.
  76. 76.Zhang, Y., Baldridge, J., and He, L. PAWS: Paraphrase Adversaries from Word Scrambling. In North American Chapter of the Association for Computational Linguistics (NAACL), 2019. URL https://aclanthology.org/N19-1131.
  77. 77.Zhou, C., Li, Q., Li, C., Yu, J., Liu, Y., Wang, G., Zhang, K., Ji, C., Yan, Q., He, L., et al. A comprehensive survey on pretrained foundation models: A history from bert to chatgpt. arXiv, 2023. URL http://arxiv.org/abs/2302.09419v3.
  78. 78.Zhou, J., Lin, Z., Zheng, Y., Li, J., and Yang, Z. Not All Tasks Are Born Equal: Understanding Zero-Shot Generalization. In International Conference on Learning Representations (ICLR), 2022.

Citation

MLA
Wang, K., et al. “Localizing Task Information for Improved Model Merging and Compression”. arXiv, 2024, http://arxiv.org/abs/2405.07813v1.
APA
Wang, K., Dimitriadis, N., Ortiz-Jimenez, G., Fleuret, F., & Frossard, P. (2024). Localizing Task Information for Improved Model Merging and Compression. arXiv. http://arxiv.org/abs/2405.07813v1
Chicago
Wang, K., N. Dimitriadis, G. Ortiz-Jimenez, F. Fleuret, and P. Frossard. 2024. “Localizing Task Information for Improved Model Merging and Compression”. arXiv. http://arxiv.org/abs/2405.07813v1.
Harvard
Wang, K. et al. (2024) “Localizing Task Information for Improved Model Merging and Compression”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2405.07813v1.
Vancouver
1. Wang K, Dimitriadis N, Ortiz-Jimenez G, Fleuret F, Frossard P (2024) Localizing Task Information for Improved Model Merging and Compression. arXiv

BibTeX

@article{wang2024localizing,
  title = {Localizing Task Information for Improved Model Merging and Compression},
  author = {Wang, Ke and Dimitriadis, Nikolaos and Ortiz-Jimenez, Guillermo and Fleuret, François and Frossard, Pascal},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2405.07813v1},
  eprint = {2405.07813}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/