Localizing Task Information for Improved Model Merging and Compression
Ke WangNikolaos DimitriadisGuillermo Ortiz-JiménezFrançois FleuretPascal Frossard
Reveals that multi-task merged models still preserve single-task information within non-overlapping parameter subsets, introducing TALL-masks and Consensus Merging to dramatically compress multiple fine-tuned checkpoints and boost merging accuracy by filtering out interfering weights.
Organizations deploying artificial intelligence increasingly rely on pre-trained foundation models fine-tuned on specialized downstream tasks. However, managing and storing dozens of massive, individual fine-tuned models incurs unsustainable computational and storage costs. Merging multiple specialized models into a single multi-task system via weight interpolation offers a potential solution, but existing approaches suffer from steep performance degradation as the number of tasks grows.
The article demonstrates that multi-task degradation is caused by task interference at evaluation rather than the erasure of task knowledge, and it presents two complementary frameworks: TALL-masks for high-ratio model compression and Consensus Merging for improved multi-task fusion.
To evaluate model behavior, the authors conducted empirical analyses across extensive vision and natural language processing benchmarks, ranging from 8 to 20 tasks using standard architectures such as Vision Transformers and T5-large. Through a controlled experiment with non-overlapping parameters, they isolated the root cause of degradation. They then constructed task-specific binary masks that isolate essential parameters from the combined multi-task weight vector and analyzed parameter agreement across tasks to filter out detrimental weights.
The evaluation revealed several key findings. First, task-relevant information is fully preserved after model merging, and applying task-specific binary masks restores over 99% of original fine-tuned accuracy across benchmarks. Second, the compression framework drastically reduces storage footprint; for instance, storing 20 fine-tuned vision models was reduced from 57.0 Gb to 8.2 Gb while retaining 99.7% of performance. Third, parameter agreement analysis identified both "catastrophic" weights (irrelevant to all tasks) and "selfish" weights (beneficial to only one task), which cause interference in combined models. Fourth, eliminating these disruptive parameters via Consensus Merging consistently improved existing fusion baselines, yielding gains such as 4.9% higher absolute accuracy on a 20-task vision benchmark and 6.3% higher accuracy on an 8-task language benchmark.
These findings indicate that teams maintaining multiple task-specific foundation models can achieve dramatic storage cost savings without sacrificing individual model accuracy. In settings requiring a single generalist model, Consensus Merging mitigates performance loss and improves scaling robustness. Practitioners can adopt Consensus Merging alongside existing pipelines since it operates directly on merged weight vectors without requiring retraining.
Organizations should evaluate their deployment requirements: apply TALL-masks when maximum individual task accuracy and aggressive storage compression are required, or apply Consensus Merging when a unified, single-checkpoint deployment is necessary. However, decision-makers should note that optimal parameter pruning thresholds depend on the underlying merging method and model fine-tuning regime. Additional validation is advised when adapting this framework to new architectures or domain-specific data.
- Paper: Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time, Mitchell Wortsman et al. (2022). Introduces weight-space averaging (Model Soups) across fine-tuned checkpoints, establishing the foundational weight-merging framework that this work analyzes and improves through localized masking.
- Paper: PackNet: Adding Multiple Tasks to a Single Network by Iterative Pruning, Arun Mallya et al. (2017). Pioneers the isolation of task capabilities within shared network parameters via binary pruning masks, providing key inspiration for TALL-masks.
- Paper: Overcoming catastrophic forgetting with hard attention to the task, Joan Serrà et al. (2018). Demonstrates how task-specific attention masks prevent cross-task interference in shared weights, establishing the prerequisite concept of parameter localization.
- Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). Examines vector-space task conflict and interference in multi-task parameter spaces, laying the groundwork for analyzing disruptive and selfish task parameters.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). Offers a comprehensive overview of parameter-efficient fine-tuning representations, which directly underpin modular task vectors used in model merging.
- Paper: EMR-Merging: Tuning-Free High-Performance Model Merging, Chenyu Huang et al. (2024). Extends the study of parameter interference in weight merging by introducing an elect, mask, and rescale paradigm that resolves task capacity limits without tuning.
