Composable Sparse Fine-Tuning for Cross-Lingual Transfer
Alan AnsellEdoardo Maria PontiAnna KorhonenIvan Vulic
Proposes Lottery Ticket Sparse Fine-Tuning, a parameter-efficient technique that learns and composes task- and language-specific parameter masks to achieve superior zero-shot cross-lingual transfer over adapter-based methods without altering model architecture or slowing inference.
Adapting large pretrained language models to new tasks across different languages is computationally expensive and frequently causes models to forget previously learned information. While modular add-on components, known as adapters, allow task and language skills to be trained separately and combined without altering base parameters, they increase the total parameter count and slow down real-time processing. Conversely, sparse fine-tuning—which updates only a fraction of existing parameters—retains the base architecture and speed but has historically lacked modularity.
The article introduces and evaluates Lottery Ticket Sparse Fine-Tuning (LT-SFT), a technique that combines the modularity of adapters with the structural efficiency of sparse fine-tuning. The main objective is to demonstrate that LT-SFT can achieve effective zero-shot cross-lingual transfer—applying a model trained on a task in one language to a new language without task-specific target data—while outperforming current adapter-based methods and preserving the underlying model's speed and architecture.
The researchers evaluated LT-SFT across four natural language processing tasks: part-of-speech tagging, dependency parsing, named entity recognition, and natural language inference. The study encompassed 35 diverse languages, focusing primarily on low-resource and underrepresented languages. The LT-SFT process works in two phases: it first performs a standard training run to identify the parameters that change the most, resets the model to its original values, and then retrains only that selected subset (typically between 1% and 8% of total parameters). These specialized sparse parameter updates for specific tasks and languages are combined with the base model via simple addition at deployment time.
The evaluation revealed several key findings. First, LT-SFT consistently outperformed the leading adapter framework (MAD-X) across all evaluated tasks, improving part-of-speech tagging accuracy by 2.5 percentage points, dependency parsing attachment score by 3.7 points, named entity recognition F1 score by 1.8 points, and natural language inference accuracy by 1.9 points. Second, unlike adapters, LT-SFT maintains the exact base architecture, preventing any loss of inference speed during deployment. Third, performance remained stable and predictable as the number of tuned parameters increased, making the method less sensitive to hyperparameter tuning than adapter frameworks. Finally, training task updates using diverse, multi-source language data substantially boosted transfer accuracy, enabling a standard base model to outperform a much larger model on cross-lingual question answering.
These results demonstrate that extreme sparsity—keeping the tuned parameter density below approximately 30%—is critical to prevent functional interference and overfitting when merging distinct task and language updates. For engineering and deployment, this approach eliminates the trade-off between modular flexibility and operational efficiency, significantly reducing computational overhead and infrastructure complexity for multilingual deployments.
Organizations deploying multilingual systems should consider adopting sparse fine-tuning as an alternative to adapter layers, particularly for low-resource languages where standard pretraining provides weak coverage. When annotated data is scarce, teams should utilize multi-source training data across diverse languages to maximize transfer performance. Future work should explore applying LT-SFT to other domains, such as multimodal applications, debiasing, and domain adaptation, while evaluating alternative parameter-selection criteria. While confidence in these empirical results is high across diverse benchmarks, stakeholders should note that performance gains are less pronounced for high-resource target languages that are already well-represented in the base pretrained models.
- Paper: Parameter-Efficient Transfer Learning for NLP, Neil Houlsby et al. (2019). It introduces the foundational modular adapter framework for parameter-efficient transfer learning that the source aims to improve upon without adding architectural parameters.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). It establishes masked language modeling and cross-lingual pre-training objectives that provide the foundational basis for language-specific transfer utilized in the source.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). It investigates zero-shot cross-lingual generalization mechanisms in multilingual pretrained models, setting up the key problem domain the source addresses.
- Paper: Overcoming catastrophic forgetting with hard attention to the task, Joan Serrà et al. (2018). It pioneers task-specific masking and gating to protect pathways and mitigate catastrophic forgetting, inspiring sparse subnet adaptation techniques.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). It provides a unified structural view of parameter-efficient transfer learning mechanisms, clarifying the design trade-offs between modular and sparse adaptation.
- Paper: AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning, Qingru Zhang et al. (2023). It extends parameter-efficient adaptation by dynamically allocating parameter budgets across model components during fine-tuning based on importance metrics.
- Paper: Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time, Mitchell Wortsman et al. (2022). It explores parameter-space weight combinations and averaging of diverse fine-tuned models without altering inference architecture, extending the broader paradigm of model composition.
- Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). It builds upon modular adaptation concepts to investigate how composing distinct parameter-efficient mechanisms mitigates catastrophic forgetting across extended horizons.
- Paper: TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents, Bofei Zhang et al. (2026). It investigates language-specific versus shared subnetwork activations in multilingual LLMs, providing mechanistic validation of modular language routing.
