Composable Sparse Fine-Tuning for Cross-Lingual Transfer

Alan AnsellEdoardo Maria PontiAnna KorhonenIvan Vulic

article2022ACL186 citations

Proposes Lottery Ticket Sparse Fine-Tuning, a parameter-efficient technique that learns and composes task- and language-specific parameter masks to achieve superior zero-shot cross-lingual transfer over adapter-based methods without altering model architecture or slowing inference.

Listen

Adapting large pretrained language models to new tasks across different languages is computationally expensive and frequently causes models to forget previously learned information. While modular add-on components, known as adapters, allow task and language skills to be trained separately and combined without altering base parameters, they increase the total parameter count and slow down real-time processing. Conversely, sparse fine-tuning—which updates only a fraction of existing parameters—retains the base architecture and speed but has historically lacked modularity.

The article introduces and evaluates Lottery Ticket Sparse Fine-Tuning (LT-SFT), a technique that combines the modularity of adapters with the structural efficiency of sparse fine-tuning. The main objective is to demonstrate that LT-SFT can achieve effective zero-shot cross-lingual transfer—applying a model trained on a task in one language to a new language without task-specific target data—while outperforming current adapter-based methods and preserving the underlying model's speed and architecture.

The researchers evaluated LT-SFT across four natural language processing tasks: part-of-speech tagging, dependency parsing, named entity recognition, and natural language inference. The study encompassed 35 diverse languages, focusing primarily on low-resource and underrepresented languages. The LT-SFT process works in two phases: it first performs a standard training run to identify the parameters that change the most, resets the model to its original values, and then retrains only that selected subset (typically between 1% and 8% of total parameters). These specialized sparse parameter updates for specific tasks and languages are combined with the base model via simple addition at deployment time.

The evaluation revealed several key findings. First, LT-SFT consistently outperformed the leading adapter framework (MAD-X) across all evaluated tasks, improving part-of-speech tagging accuracy by 2.5 percentage points, dependency parsing attachment score by 3.7 points, named entity recognition F1 score by 1.8 points, and natural language inference accuracy by 1.9 points. Second, unlike adapters, LT-SFT maintains the exact base architecture, preventing any loss of inference speed during deployment. Third, performance remained stable and predictable as the number of tuned parameters increased, making the method less sensitive to hyperparameter tuning than adapter frameworks. Finally, training task updates using diverse, multi-source language data substantially boosted transfer accuracy, enabling a standard base model to outperform a much larger model on cross-lingual question answering.

These results demonstrate that extreme sparsity—keeping the tuned parameter density below approximately 30%—is critical to prevent functional interference and overfitting when merging distinct task and language updates. For engineering and deployment, this approach eliminates the trade-off between modular flexibility and operational efficiency, significantly reducing computational overhead and infrastructure complexity for multilingual deployments.

Organizations deploying multilingual systems should consider adopting sparse fine-tuning as an alternative to adapter layers, particularly for low-resource languages where standard pretraining provides weak coverage. When annotated data is scarce, teams should utilize multi-source training data across diverse languages to maximize transfer performance. Future work should explore applying LT-SFT to other domains, such as multimodal applications, debiasing, and domain adaptation, while evaluating alternative parameter-selection criteria. While confidence in these empirical results is high across diverse benchmarks, stakeholders should note that performance gains are less pronounced for high-resource target languages that are already well-represented in the base pretrained models.

Cover for Composable Sparse Fine-Tuning for Cross-Lingual Transfer

Abstract

Fine-tuning the entire set of parameters of a large pretrained model has become the mainstream approach for transfer learning. To increase its efficiency and prevent catastrophic forgetting and interference, techniques like adapters and sparse fine-tuning have been developed. Adapters are modular, as they can be combined to adapt a model towards different facets of knowledge (e.g., dedicated language and/or task adapters). Sparse fine-tuning is expressive, as it controls the behavior of all model components. In this work, we introduce a new fine-tuning method with both these desirable properties. In particular, we learn sparse, real-valued masks based on a simple variant of the Lottery Ticket Hypothesis. Task-specific masks are obtained from annotated data in a source language, and language-specific masks from masked language modeling in a target language. Both these masks can then be composed with the pretrained model. Unlike adapter-based fine-tuning, this method neither increases the number of parameters at inference time nor alters the original model architecture. Most importantly, it outperforms adapters in zero-shot cross-lingual transfer by a large margin in a series of multilingual benchmarks, including Universal Dependencies, MasakhaNER, and AmericasNLI. Based on an in-depth analysis, we additionally find that sparsity is crucial to prevent both 1) interference between the fine-tunings to be composed and 2) overfitting. We release the code and models at https://github.com/cambridgeltl/composable-sft.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Methodology
  • 3.1 Lottery Ticket Sparse Fine-Tuning
  • 3.2 Zero-Shot Transfer with LT-SFT
  • 4 Experimental Setup
  • 4.1 Baselines and Model Variants
  • 4.2 Language SFT/Adapter Training Setup
  • 4.3 Task SFT/Adapter Training Setup
  • 4.4 Multi-Source Training
  • 5 Results and Discussion
  • 5.1 Multi-Source Training
  • 5.2 Benefits of Sparsity
  • 6 Related Work
  • 7 Conclusion and Future Work
  • Acknowledgements
  • References
  • A Algorithm of Cross-Lingual Transfer with LT-SFT
  • B Languages
  • C Results by Language
  • D MAD-X Results with AdapterHub Adapters
  • E Parameter Overlap between Languages

Knowls

  1. Knowl 1 — Lottery Ticket Sparse Fine-Tuning

    model/method

    Lottery Ticket Sparse Fine-Tuning (LT-SFT) adapts a pretrained neural model F(⋅;θ(0))F(\cdot; \theta^{(0)}) with parameter vector heta(0)∈RN heta^{(0)} \in \mathbb{R}^N by identifying and fine-tuning a task- or language-specific sparse difference vector ϕ∈RN\phi \in \mathbb{R}^N. The training procedure operates in two successive phases:

    1. Phase 1 (Dense Exploration and Mask Selection): The base parameters θ(0)\theta^{(0)} are fully fine-tuned on target data D\mathcal{D} using loss function L\mathcal{L} until convergence, producing parameter state θ(1)\theta^{(1)}. Model parameters are ranked by their absolute deviation from the pretrained weights, ∣θi(1)−θi(0)∣|\theta^{(1)}_i - \theta^{(0)}_i|, and a binary mask μ∈{0,1}N\mu \in \{0, 1\}^N is constructed to select the top KK parameters with the largest change: μi={1if ∣θi(1)−θi(0)∣ is in the top K0otherwise\mu_i = \begin{cases} 1 & \text{if } |\theta^{(1)}_i - \theta^{(0)}_i| \text{ is in the top } K \\ 0 & \text{otherwise} \end{cases}

    2. Phase 2 (Sparse Optimization): Parameters are reset to their original pretrained values θ(0)\theta^{(0)}. The model is fine-tuned on D\mathcal{D} updating only the KK selected parameters by applying the masked gradient μ⊙∇θL(F(⋅;θ),D)\mu \odot \nabla_\theta \mathcal{L}(F(\cdot; \theta), \mathcal{D}) at each optimization step, while the remaining N−KN - K parameters remain frozen at θ(0)\theta^{(0)}.

    Upon convergence to parameter state θ(2)\theta^{(2)}, the resulting sparse difference vector is computed as ϕ=θ(2)−θ(0)\phi = \theta^{(2)} - \theta^{(0)}. To discourage parameters from drifting excessively from initialization during language adaptation, an L1L_1 regularization term is added to the objective: J(θ)=λN∑i=1N∣θi−θi(0)∣J(\theta) = \frac{\lambda}{N} \sum_{i=1}^N |\theta_i - \theta^{(0)}_i| with regularization strength λ=0.1\lambda = 0.1.

  2. Knowl 2 — Zero-Shot Cross-Lingual Transfer with Composable Sparse Fine-Tunings

    model/method

    Zero-shot cross-lingual transfer is achieved by linearly composing language-specific and task-specific sparse fine-tunings (SFTs) onto a shared pretrained massively multilingual Transformer F(⋅;θ(0))F(\cdot; \theta^{(0)}):

    1. Language Adaptation: For each target language ll, an SFT difference vector ϕL(l)\phi^{(l)}_L is learned using masked language modeling (MLM) via LT-SFT on monolingual text of language ll.
    2. Task Adaptation with Source Decoupling: Given task tt and labeled data in source language ss (e.g., English), the model is first initialized with the source language SFT: θ=θ(0)+ϕL(s)\theta = \theta^{(0)} + \phi^{(s)}_L. A task classifier head is added, and both the model's top KK parameters and the classifier head are fine-tuned with LT-SFT. Once training yields converged model parameters θ′\theta', the source language difference is subtracted to isolate the pure task difference vector: ϕT(t)=θ′−(θ(0)+ϕL(s))\phi^{(t)}_T = \theta' - \left(\theta^{(0)} + \phi^{(s)}_L\right)
    3. Zero-Shot Target Composition: For target language ll and task tt, the task and target language difference vectors are directly added to the base pretrained model parameters: Ft,l=F(⋅;θ(0)+ϕT(t)+ϕL(l))F_{t, l} = F\left(\cdot; \theta^{(0)} + \phi^{(t)}_T + \phi^{(l)}_L\right) The trained task classifier head is placed on top of Ft,lF_{t, l} for downstream inference without modifying the underlying architecture or increasing inference parameter count.
  3. Knowl 3 — Algorithm for Zero-Shot Cross-Lingual Transfer with LT-SFT

    algorithm

    The algorithm formalizes the generation of sparse difference vectors via two-phase lottery ticket selection and their linear composition for cross-lingual zero-shot inference.

    function LTSFT(D, L, \theta^{(0)}, \eta, K)
        \theta^{(1)} \leftarrow \theta^{(0)}
        while not converged do
            \theta^{(1)} \leftarrow \theta^{(1)} - \eta \nabla \mathcal{L}(\theta^{(1)}, D)
        for each parameter index i in 1 to N do
            if |\theta^{(1)}_i - \theta^{(0)}_i| is in top K then
                \mu_i \leftarrow 1
            else
                \mu_i \leftarrow 0
        \theta^{(2)} \leftarrow \theta^{(0)}
        while not converged do
            \theta^{(2)} \leftarrow \theta^{(2)} - \mu \odot \eta \nabla \mathcal{L}(\theta^{(2)}, D)
        \phi \leftarrow \theta^{(2)} - \theta^{(0)}
        return \phi
    function CROSSLINGUALTRANSFER(D_{src}, D_{tar}, D_{task}, \mathcal{L}_{task}, \theta^{(0)}, \eta, K)
        \phi_{src} \leftarrow LTSFT(D_{src}, \mathcal{L}_{MLM}, \theta^{(0)}, \eta, K)
        \phi_{task} \leftarrow LTSFT(D_{task}, \mathcal{L}_{task}, \theta^{(0)} + \phi_{src}, \eta, K)
        \phi_{tar} \leftarrow LTSFT(D_{tar}, \mathcal{L}_{MLM}, \theta^{(0)}, \eta, K)
        return \theta^{(0)} + \phi_{task} + \phi_{tar}

    Optimization details:

    • Optimizer: AdamW with initial learning rate η=5×10−5\eta = 5 \times 10^{-5} linearly decayed to 00 (2×10−52 \times 10^{-5} for NLI).
    • Language SFT training: min⁡(100 epochs,100,000 steps)\min(100 \text{ epochs}, 100{,}000 \text{ steps}) with batch size 8, sequence length 256, minimum 30,000 steps.
    • Task SFT training: Phase 1 runs for 3 epochs (to prevent early overfitting), Phase 2 runs for 10 epochs (5 epochs for NLI).
  4. Knowl 4 — Zero-Shot Cross-Lingual Transfer Performance Across Multilingual Benchmarks

    data/table

    Zero-shot cross-lingual transfer was evaluated across 35 typologically diverse languages on four tasks: Part-of-Speech Tagging (POS; Universal Dependencies 2.7, accuracy), Dependency Parsing (DP; Universal Dependencies 2.7, UAS and LAS), Named Entity Recognition (NER; MasakhaNER, F1), and Natural Language Inference (NLI; AmericasNLI, accuracy). Base multilingual models are mBERT (for POS, DP, NER) and XLM-R (for NLI). English is the source language.

    Method POS Acc DP UAS DP LAS NER F1 NLI Acc
    LT-SFT 71.1 (1) 57.1 (1) 37.8 (1) 71.7 (1) 51.4 (1)
    RAND-SFT 69.2 (1) 54.3 (1) 33.9 (1) - -
    MAD-X 68.6 (16) 54.6 (2) 34.1 (1) 69.9 (8) 49.5 (2)
    BITFIT 58.1 45.7 23.9 54.9 38.3
    LT-SFT TA-ONLY 51.3 (32) 39.1 (1) 19.9 (1) 55.3 (8) 39.9 (4)
    MAD-X TA-ONLY 52.1 (32) 38.9 (1) 19.5 (1) 52.4 (32) 41.7 (4)

    (Numbers in parentheses represent the optimal adapter reduction factor or equivalent SFT parameter budget KK, where factor 1 corresponds to ~8.0% sparsity on mBERT and ~5.1% on XLM-R).

    LT-SFT consistently outperforms the MAD-X adapter baseline (+2.5 POS accuracy, +2.5 DP UAS, +3.7 DP LAS, +1.8 NER F1, +1.9 NLI accuracy) and random parameter selection (RAND-SFT). Language adaptation provides substantial improvements over task-only adaptation (TA-ONLY) on unseen low-resource languages.

  5. Knowl 5 — Decoupling and Output Embedding Freezing in Language SFT

    model/method

    When performing language sparse fine-tuning with masked language modeling (MLM), standard unconstrained Phase 1 fine-tuning causes the overwhelming majority of the top KK parameter updates to fall into the output token embedding matrix due to its direct adjacency to the prediction loss. This parameter concentration degrades downstream representation transfer.

    To ensure balanced parameter selection across deeper representation layers:

    1. Input and output token embedding matrices are decoupled.
    2. The parameters of the output embedding matrix are kept strictly frozen throughout language adaptation.
    3. Layer normalization parameters are also frozen. All other model weights (input embeddings, multi-head self-attention, and intermediate feed-forward projections) remain trainable during Phase 1 ranking and Phase 2 sparse optimization.
  6. Knowl 6 — Multi-Source Task Training with Composable SFTs

    empirical result

    Task SFTs can be trained on pooled annotated data across multiple source languages by alternating language-specific SFT difference vectors per training batch: each batch contains examples from a single source language ss, and the corresponding pretrained language SFT ϕL(s)\phi^{(s)}_L is activated during that forward/backward step.

    Multi-source LT-SFT delivers substantial gains over single-source (English) transfer:

    1. Dependency Parsing (Universal Dependencies): Multi-source training over treebanks from 11 high-resource languages boosts zero-shot transfer across target languages from 57.1 to 64.3 UAS and from 37.8 to 47.6 LAS.
    2. Natural Language Inference (AmericasNLI): Training on MultiNLI plus 14 XNLI languages raises average accuracy from 51.4% to 53.1%.
    3. Extractive Question Answering (XQuAD evaluation on non-MLQA languages el, ro, ru, th, tr): XLM-R Base trained with multi-source LT-SFT achieves F1/EM scores of 81.9/65.5 (Greek), 86.3/73.3 (Romanian), 81.4/64.6 (Russian), 82.4/75.2 (Thai), and 75.2/58.6 (Turkish). This configuration outperforms full fine-tuning of XLM-R Base (71.1/54.3, 78.3/63.7, 74.1/57.8, 67.1/55.7, 67.5/51.1) and surpasses the larger XLM-R Large model under full fine-tuning (79.8/61.7, 83.6/69.7, 80.1/64.3, 74.2/62.8, 75.9/59.3).
  7. Knowl 7 — Impact of Density on SFT Compositionality and Knowledge Interference

    empirical result

    Systematic sweeps across task and language SFT parameter densities (the percentage of non-zero parameters tuned, from 5% to 100%) on Dependency Parsing and Named Entity Recognition show that sparsity is necessary for successful linear composition:

    • Zero-shot transfer performance is highest and remains stable when both task and language SFTs are sparse.
    • When the density of tuned parameters exceeds ~30%, compositional transfer performance degrades significantly across tasks.
    • At higher densities, overlapping parameter modifications between task and language vectors create negative interference, and unconstrained parameter capacity increases susceptibility to overfitting.
    • Task fine-tuning performance remains flat above ~60% density because unused vocabulary token embeddings in the embedding layer receive no gradient updates during task training.
  8. Knowl 8 — Robustness of LT-SFT to Tunable Parameter Budget Scaling

    empirical result

    Evaluating task adaptation capacity across adapter reduction factors (32, 16, 8, 4, 2, 1) and equivalent LT-SFT parameter counts KK (ranging from 0.25% to 8.0% sparsity on mBERT, corresponding to ~442K to 14.2M trainable parameters) highlights key scaling differences:

    • LT-SFT Scaling: Performance monotonically improves or stabilizes as KK increases, with peak zero-shot accuracy consistently occurring at an equivalent reduction factor of 1 across all four evaluated tasks (POS, DP, NER, NLI).
    • MAD-X Scaling: Adapter performance does not scale monotonically with capacity; lower reduction factors (larger adapters) frequently cause performance degradation, making the optimal reduction factor task-dependent (factor 16 for POS, factor 2 for DP, factor 8 for NER, and factor 2 for NLI).
    • Inference Efficiency: Increasing KK in LT-SFT has zero impact on inference latency or model parameter count because ϕ\phi is merged into base weights before inference, whereas increasing adapter capacity in MAD-X incurs additional computational overhead.
  9. Knowl 9 — Language Adaptation Disparity Between Pretrained-Seen and Pretrained-Unseen Languages

    empirical result

    The downstream benefit of language-specific adaptation depends strongly on whether the language was observed during multilingual pretraining and its pretraining resource scale:

    • Low-Resource Unseen Languages: Language SFT adaptation provides large performance gains over task-only adaptation (e.g., POS accuracy improves from 51.3% to 71.1% on average across unseen languages).
    • High-Resource Seen Languages: For languages well-represented in mBERT pretraining (Arabic, Japanese, Chinese), target language adaptation provides negligible benefit over task-only adaptation (POS accuracy is 63.4% with LT-SFT vs. 63.5% with LT-SFT TA-ONLY; DP UAS is 55.9% vs. 54.4%).
    • Low-Resource Seen Languages: For underrepresented seen languages (Swahili, Yorùbá on NER), language SFT adaptation yields clear improvements (average F1 of 77.1% vs. 69.4% for TA-ONLY), showing that lower-resource languages have greater headroom for specialization via MLM SFT.
  10. Knowl 10 — Sub-Network Parameter Overlap Across Diverse Languages

    empirical result

    Measuring the pairwise parameter overlap percentage between language SFT masks selected by LT-SFT across 25 languages reveals that overlap is universally low (ranging between 8% and 12%):

    • Except for closely related languages sharing orthography and vocabulary (such as Mandarin Chinese and Cantonese), typologically distinct languages select largely disjoint sub-networks of parameters.
    • Because multilingual Transformers contain multiple alternative winning tickets that achieve comparable adaptation performance, LT-SFT identifies specialized, non-interfering sub-networks for individual languages within the shared parameter space.

Coverage note — No substantial contributed material was omitted. All key algorithmic mechanisms, experimental configurations, main benchmark transfer results, multi-source training findings, sparsity analyses, and language overlap studies are covered.

References

  1. 1.David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Anuoluwapo Aremu, Catherine Gitau, Derguene Mbaye, Jesujoba Alabi, Seid Muhie Yimam, Tajuddeen Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin Adewumi, Paul Rayson, Mofetoluwa Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane MBOUP, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima DIOP, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, and Salomey Osei. 2021. MasakhaNER: Named Entity Recognition for African Languages. arXiv preprint.
  2. 2.Željko Agić and Ivan Vulić. 2019. JW300: A wide-coverage parallel corpus for low-resource languages. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3204–3210, Florence, Italy. Association for Computational Linguistics.
  3. 3.Alan Ansell, Edoardo Maria Ponti, Jonas Pfeiffer, Sebastian Ruder, Goran Glavaš, Ivan Vulić, and Anna Korhonen. 2021. MAD-G: Multilingual adapter generation for efficient cross-lingual transfer. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4762–4781, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  4. 4.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637, Online. Association for Computational Linguistics.
  5. 5.Ankur Bapna and Orhan Firat. 2019. Simple, scalable adaptation for neural machine translation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1538–1548, Hong Kong, China. Association for Computational Linguistics.
  6. 6.David Brambila. 1976. Diccionario Raramuri-Castellano: Tarahumar.
  7. 7.Gina Bustamante, Arturo Oncevay, and Roberto Zariquiey. 2020. No data to crawl? monolingual corpus creation from PDF files of truly low-resource languages in Peru. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 2914–2923, Marseille, France. European Language Resources Association.
  8. 8.Luis Chiruzzo, Pedro Amarilla, Adolfo Ríos, and Gustavo Giménez Lugo. 2020. Development of a Guarani - Spanish parallel corpus. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 2629–2633, Marseille, France. European Language Resources Association.
  9. 9.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  10. 10.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  11. 11.Rubén Cushimariano Romano and Richer C. Sebastián Q. 2008. Ñaantsipeta asháninkaki birakochaki. diccionario asháninka-castellano. versión preliminar. http://www.lengamer.org/publicaciones/diccionarios/.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Timothy Dozat and Christopher D. Manning. 2017. Deep biaffine attention for neural dependency parsing. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  14. 14.Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir, Gustavo A. Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando A. Coto Solano, Ngoc Thang Vu, and Katharina Kann. 2021. AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource Languages.
  15. 15.Isaac Feldman and Rolando Coto-Solano. 2020. Neural machine translation models with back-translation for the extremely low-resource indigenous language Bribri. In Proceedings of the 28th International Conference on Computational Linguistics, pages 3965–3976, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  16. 16.Jonathan Frankle and Michael Carbin. 2019. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In International Conference on Learning Representations.
  17. 17.Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M Roy, and Michael Carbin. 2019. Stabilizing the lottery ticket hypothesis. arXiv preprint arXiv:1903.01611.
  18. 18.Ana-Paula Galarreta, Andrés Melgar, and Arturo Oncevay. 2017. Corpus creation and initial SMT experiments between Spanish and Shipibo-konibo. In Proceedings of the International Conference Recent Advances in Natural Language Processing, RANLP 2017, pages 238–244, Varna, Bulgaria. INCOMA Ltd.
  19. 19.Goran Glavaš and Ivan Vulić. 2021. Is supervised syntactic parsing beneficial for language understanding tasks? an empirical investigation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 3090–3104, Online. Association for Computational Linguistics.
  20. 20.Demi Guo, Alexander Rush, and Yoon Kim. 2021. Parameter-efficient transfer learning with diff pruning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4884–4896, Online. Association for Computational Linguistics.
  21. 21.Ximena Gutierrez-Vasques, Gerardo Sierra, and Isaac Hernandez Pompa. 2016. Axolotl: a web accessible parallel corpus for Spanish-Nahuatl. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 4210–4214, Portorož, Slovenia. European Language Resources Association (ELRA).
  22. 22.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-Efficient Transfer Learning for NLP. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 2790–2799. PMLR.
  23. 23.Jeremy Howard and Sebastian Ruder. 2018. Universal language model fine-tuning for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 328–339, Melbourne, Australia. Association for Computational Linguistics.
  24. 24.Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020. MLQA: Evaluating cross-lingual extractive question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7315–7330, Online. Association for Computational Linguistics.
  25. 25.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations.
  26. 26.Manuel Mager, Diónico Carrillo, and Ivan Meza. 2018. Probabilistic finite-state morphological segmenter for wixarika (huichol) language. Journal of Intelligent & Fuzzy Systems, 34(5):3081–3087.
  27. 27.Eran Malach, Gilad Yehudai, Shai Shalev-Schwartz, and Ohad Shamir. 2020. Proving the lottery ticket hypothesis: Pruning is all you need. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 6682–6691. PMLR.
  28. 28.Elena Mihas. 2011. Añaani katonkosatzi parenini, El idioma del alto Perené. Milwaukee, WI: Clarks Graphics.
  29. 29.John E Ortega, Richard Alexander Castro-Mamani, and Jaime Rafael Montoya Samame. 2020. Overcoming resistance: The normalization of an Amazonian tribal language. In Proceedings of the 3rd Workshop on Technologies for MT of Low Resource Languages, pages 1–13, Suzhou, China. Association for Computational Linguistics.
  30. 30.Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2021a. AdapterFusion: Non-destructive task composition for transfer learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 487–503, Online. Association for Computational Linguistics.
  31. 31.Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulić, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020a. AdapterHub: A framework for adapting transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 46–54, Online. Association for Computational Linguistics.
  32. 32.Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2020b. MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7654–7673, Online. Association for Computational Linguistics.
  33. 33.Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2021b. UNKs everywhere: Adapting multilingual language models to new scripts. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  34. 34.Edoardo Ponti, Ivan Vulić, Ryan Cotterell, Marinela Parovic, Roi Reichart, and Anna Korhonen. 2021. Parameter space factorization for zero-shot learning across tasks and languages. Transactions of the Association for Computational Linguistics, 9(0):410–428.
  35. 35.Edoardo Maria Ponti. 2021. Inductive Bias and Modular Design for Sample-Efficient Neural Language Learning. Ph.D. thesis, University of Cambridge.
  36. 36.Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020. XCOPA: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2362–2376, Online. Association for Computational Linguistics.
  37. 37.Edoardo Maria Ponti, Helen O’Horan, Yevgeni Berzak, Ivan Vulić, Roi Reichart, Thierry Poibeau, Ekaterina Shutova, and Anna Korhonen. 2019. Modeling language variation and universals: A survey on typological linguistics for natural language processing. Computational Linguistics, 45(3):559–601.
  38. 38.Sai Prasanna, Anna Rogers, and Anna Rumshisky. 2020. When BERT Plays the Lottery, All Tickets Are Winning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3208–3229, Online. Association for Computational Linguistics.
  39. 39.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  40. 40.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017. Learning Multiple Visual Domains with Residual Adapters. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  41. 41.Alex Renda, Jonathan Frankle, and Michael Carbin. 2020. Comparing rewinding and fine-tuning in neural network pruning. In International Conference on Learning Representations.
  42. 42.Jörg Tiedemann. 2012. Parallel data, tools and interfaces in OPUS. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12), pages 2214–2218, Istanbul, Turkey. European Language Resources Association (ELRA).
  43. 43.Erik F. Tjong Kim Sang and Fien De Meulder. 2003. Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, pages 142–147.
  44. 44.Ahmet Üstün, Arianna Bisazza, Gosse Bouma, and Gertjan van Noord. 2020. UDapter: Language adaptation for truly Universal Dependency parsing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2302–2315.
  45. 45.Marko Vidoni, Ivan Vulić, and Goran Glavaš. 2020. Orthogonal language and task adapters in zero-shot cross-lingual transfer. CoRR, abs/2012.06460.
  46. 46.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations.
  47. 47.Zirui Wang, Zachary C. Lipton, and Yulia Tsvetkov. 2020. On negative interference in multilingual models: Findings and a meta-learning treatment. In Proceedings of EMNLP 2020, pages 4438–4450.
  48. 48.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  49. 49.Mitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi, Mohammad Rastegari, Jason Yosinski, and Ali Farhadi. 2020. Supermasks in superposition. In Advances in Neural Information Processing Systems, volume 33, pages 15173–15184. Curran Associates, Inc.
  50. 50.Dongkuan Xu, Ian En-Hsu Yen, Jinxi Zhao, and Zhibin Xiao. 2021a. Rethinking network pruning – under the pre-train and fine-tune paradigm. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2376–2382, Online. Association for Computational Linguistics.
  51. 51.Runxin Xu, Fuli Luo, Zhiyuan Zhang, Chuanqi Tan, Baobao Chang, Songfang Huang, and Fei Huang. 2021b. Raise a child in large language model: Towards effective and generalizable fine-tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9514–9528, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  52. 52.Haonan Yu, Sergey Edunov, Yuandong Tian, and Ari S. Morcos. 2020. Playing the lottery with rewards and multiple languages: lottery tickets in rl and nlp. In International Conference on Learning Representations.
  53. 53.Elad Ben Zaken, Shauli Ravfogel, and Yoav Goldberg. 2021. BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models. CoRR, abs/2106.10199.
  54. 54.Daniel Zeman, Joakim Nivre, Mitchell Abrams, Elia Ackermann, Noëmi Aepli, Hamid Aghaei, Željko Agić, Amir Ahmadi, Lars Ahrenberg, Chika Kennedy Ajede, Gabrielė Aleksandravičiūtė, Ika Alfina, Lene Antonsen, Katya Aplonova, Angelina Aquino, Carolina Aragon, Maria Jesus Aranzabe, Hórunn Arnardóttir, Gashaw Arutie, Jessica Naraiswari Arwidarasti, Masayuki Asahara, Luma Ateyah, Furkan Atmaca, Mohammed Attia, Aitziber Atutxa, Liesbeth Augustinus, Elena Badmaeva, Keerthana Balasubramani, Miguel Ballesteros, Esha Banerjee, Sebastian Bank, Verginica Barbu Mititelu, Victoria Basmov, Colin Batchelor, John Bauer, Seyyit Talha Bedir, Kepa Bengoetxea, Gözde Berk, Yevgeni Berzak, Irshad Ahmad Bhat, Riyaz Ahmad Bhat, Erica Biagetti, Eckhard Bick, Agnė Bielinskienė, Kristín Bjarnadóttir, Rogier Blokland, Victoria Bobicev, Loïc Boizou, Emanuel Borges Völker, Carl Börstell, Cristina Bosco, Gosse Bouma, Sam Bowman, Adriane Boyd, Kristina Brokaitė, Aljoscha Burchardt, Marie Candito, Bernard Caron, Gauthier Caron, Tatiana Cavalcanti, Gülşen Cebiroğlu Eryiğit, Flavio Massimiliano Cecchini, Giuseppe G. A. Celano, Slavomír Céplö, Savas Cetin, Özlem Çetinoğlu, Fabricio Chalub, Ethan Chi, Yongseok Cho, Jinho Choi, Jayeol Chun, Alessandra T. Cignarella, Silvie Cinková, Aurélie Collomb, Çağrı Çöltekin, Miriam Connor, Marine Courtin, Elizabeth Davidson, Marie-Catherine de Marneffe, Valeria de Paiva, Mehmet Oguz Derin, Elvis de Souza, Arantza Diaz de Ilarraza, Carly Dickerson, Arawinda Dinakaramani, Bamba Dione, Peter Dirix, Kaja Dobrovoljc, Timothy Dozat, Kira Droganova, Puneet Dwivedi, Hanne Eckhoff, Marhaba Eli, Ali Elkahky, Binyam Ephrem, Olga Erina, Tomaž Erjavec, Aline Etienne, Wograine Evelyn, Sidney Facundes, Richárd Farkas, Marília Fernanda, Hector Fernandez Alcalde, Jennifer Foster, Cláudia Freitas, Kazunori Fujita, Katarína Gajdošová, Daniel Galbraith, Marcos Garcia, Moa Gärdenfors, Sebastian Garza, Fabrício Ferraz Gerardi, Kim Gerdes, Filip Ginter, Iakes Goenaga, Koldo Gojenola, Memduh Gökırmak, Yoav Goldberg, Xavier Gómez Guinovart, Berta González Saavedra, Bernadeta Griciūtė, Matias Grioni, Loïc Grobol, Normunds Grūzītis, Bruno Guillaume, Céline Guillot-Barbance, Tunga Güngör, Nizar Habash, Hinrik Hafsteinsson, Jan Hajič, Jan Hajič jr., Mika Hämäläinen, Linh Hà Mỹ, Na-Rae Han, Muhammad Yudistira Hanifmuti, Sam Hardwick, Kim Harris, Dag Haug, Johannes Heinecke, Oliver Hellwig, Felix Hennig, Barbora Hladká, Jaroslava Hlaváčová, Florinel Hociung, Petter Hohle, Eva Huber, Jena Hwang, Takumi Ikeda, Anton Karl Ingason, Radu Ion, Elena Irimia, Ọlájídé Ishola, Tomáš Jelínek, Anders Johannsen, Hildur Jónsdóttir, Fredrik Jørgensen, Markus Juutinen, Sarveswaran K, Hüner Kaşıkara, Andre Kaasen, Nadezhda Kabaeva, Sylvain Kahane, Hiroshi Kanayama, Jenna Kanerva, Boris Katz, Tolga Kayadelen, Jessica Kenney, Václava Kettnerová, Jesse Kirchner, Elena Klementieva, Arne Köhn, Abdullatif Köksal, Kamil Kopacewicz, Timo Korkiakangas, Natalia Kotsyba, Jolanta Kovalevskaitė, Simon Krek, Parameswari Krishnamurthy, Sookyoung Kwak, Veronika Laippala, Lucia Lam, Lorenzo Lambertino, Tatiana Lando, Septina Dian Larasati, Alexei Lavrentiev, John Lee, Phương Lê H`ông, Alessandro Lenci, Saran Lertpradit, Herman Leung, Maria Levina, Cheuk Ying Li, Josie Li, Keying Li, Yuan Li, KyungTae Lim, Krister Lindén, Nikola Ljubešić, Olga Loginova, Andry Luthfi, Mikko Luukko, Olga Lyashevskaya, Teresa Lynn, Vivien Macketanz, Aibek Makazhanov, Michael Mandl, Christopher Manning, Ruli Manurung, Cătălina Mărănduc, David Mareček, Katrin Marheinecke, Héctor Martínez Alonso, André Martins, Jan Mašek, Hiroshi Matsuda, Yuji Matsumoto, Ryan McDonald, Sarah McGuinness, Gustavo Mendonça, Niko Miekka, Karina Mischenkova, Margarita Misirpashayeva, Anna Missilä, Cătălin Mititelu, Maria Mitrofan, Yusuke Miyao, AmirHossein Mojiri Foroushani, Amirsaeid Moloodi, Simonetta Montemagni, Amir More, Laura Moreno Romero, Keiko Sophie Mori, Shinsuke Mori, Tomohiko Morioka, Shigeki Moro, Bjartur Mortensen, Bohdan Moskalevskyi, Kadri Muischnek, Robert Munro, Yugo Murawaki, Kaili Müürisep, Pinkey Nainwani, Mariam Nakhlé, Juan Ignacio Navarro Horñiacek, Anna Nedoluzhko, Gunta Nešpore-Bērzkalne, Lương Nguyễn Thị, Huyền Nguyễn Thị Minh, Yoshihiro Nikaido, Vitaly Nikolaev, Rattima Nitisaroj, Alireza Nourian, Hanna Nurmi, Stina Ojala, Atul Kr. Ojha, Adédayọ̀ Olúòkun, Mai Omura, Emeka Onwuegbuzia, Petya Osenova, Robert Östling, Lilja Øvrelid, Şaziye Betül Özateş, Arzucan Özgür, Balkız Öztürk Başaran, Niko Partanen, Elena Pascual, Marco Passarotti, Agnieszka Patejuk, Guilherme Paulino-Passos, Angelika Peljak-Łapińska, Siyao Peng, Cenel-Augusto Perez, Natalia Perkova, Guy Perrier, Slav Petrov, Daria Petrova, Jason Phelan, Jussi Piitulainen, Tommi A Pirinen, Emily Pitler, Barbara Plank, Thierry Poibeau, Larisa Ponomareva, Martin Popel, Lauma Pretkalniņa, Sophie Prévost, Prokopis Prokopidis, Adam Przepiórkowski, Tiina Puolakainen, Sampo Pyysalo, Peng Qi, Andriela Rääbis, Alexandre Rademaker, Taraka Rama, Loganathan Ramasamy, Carlos Ramisch, Fam Rashel, Mohammad Sadegh Rasooli, Vinit Ravishankar, Livy Real, Petru Rebeja, Siva Reddy, Georg Rehm, Ivan Riabov, Michael Rießler, Erika Rimkutė, Larissa Rinaldi, Laura Rituma, Luisa Rocha, Eiríkur Rögnvaldsson, Mykhailo Romanenko, Rudolf Rosa, Valentin Roșca, Davide Rovati, Olga Rudina, Jack Rueter, Kristján Rúnarsson, Shoval Sadde, Pegah Safari, Benoît Sagot, Aleksi Sahala, Shadi Saleh, Alessio Salomoni, Tanja Samardžić, Stephanie Samson, Manuela Sanguinetti, Dage Särg, Baiba Saulīte, Yanin Sawanakunanon, Kevin Scannell, Salvatore Scarlata, Nathan Schneider, Sebastian Schuster, Djamé Seddah, Wolfgang Seeker, Mojgan Seraji, Mo Shen, Atsuko Shimada, Hiroyuki Shirasu, Muh Shohibussirri, Dmitry Sichinava, Einar Freyr Sigurðsson, Aline Silveira, Natalia Silveira, Maria Simi, Radu Simionescu, Katalin Simkó, Mária Šimková, Kiril Simov, Maria Skachedubova, Aaron Smith, Isabela Soares-Bastos, Carolyn Spadine, Steinthór Steingrímsson, Antonio Stella, Milan Straka, Emmett Strickland, Jana Strnadová, Alane Suhr, Yogi Lesmana Sulestio, Umut Sulubacak, Shingo Suzuki, Zsolt Szántó, Dima Taji, Yuta Takahashi, Fabio Tamburini, Mary Ann C. Tan, Takaaki Tanaka, Samson Tella, Isabelle Tellier, Guillaume Thomas, Liisi Torga, Marsida Toska, Trond Trosterud, Anna Trukhina, Reut Tsarfaty, Utku Türk, Francis Tyers, Sumire Uematsu, Roman Untilov, Zdenka Urešová, Larraitz Uria, Hans Uszkoreit, Andrius Utka, Sowmya Vajjala, Daniel van Niekerk, Gertjan van Noord, Viktor Varga, Eric Villemonte de la Clergerie, Veronika Vincze, Aya Wakasa, Joel C. Wallenberg, Lars Wallin, Abigail Walsh, Jing Xian Wang, Jonathan North Washington, Maximilan Wendt, Paul Widmer, Seyi Williams, Mats Wirén, Christian Wittern, Tsegay Woldemariam, Tak-sum Wong, Alina Wróblewska, Mary Yako, Kayo Yamashita, Naoki Yamazaki, Chunxiao Yan, Koichi Yasuoka, Marat M. Yavrumyan, Zhuoran Yu, Zdeněk Žabokrtský, Shorouq Zahra, Amir Zeldes, Hanzhi Zhu, and Anna Zhuravleva. 2020. Universal Dependencies 2.7. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University.
  55. 55.Hattie Zhou, Janice Lan, Rosanne Liu, and Jason Yosinski. 2019. Deconstructing lottery tickets: Zeros, signs, and the supermask. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.

Citation

MLA
Ansell, A., et al. “Composable Sparse Fine-Tuning for Cross-Lingual Transfer”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 1778–96, https://doi.org/10.18653/v1/2022.acl-long.125.
APA
Ansell, A., Ponti, E. M., Korhonen, A., & Vulić, I. (2022). Composable Sparse Fine-Tuning for Cross-Lingual Transfer. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1778–1796. https://doi.org/10.18653/v1/2022.acl-long.125
Chicago
Ansell, A., E. M. Ponti, A. Korhonen, and I. Vulić. 2022. “Composable Sparse Fine-Tuning for Cross-Lingual Transfer”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1778–96. https://doi.org/10.18653/v1/2022.acl-long.125.
Harvard
Ansell, A. et al. (2022) “Composable Sparse Fine-Tuning for Cross-Lingual Transfer”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1778–1796. Available at: https://doi.org/10.18653/v1/2022.acl-long.125.
Vancouver
1. Ansell A, Ponti EM, Korhonen A, Vulić I (2022) Composable Sparse Fine-Tuning for Cross-Lingual Transfer. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1778–1796

BibTeX

@inproceedings{ansell-etal-2022-composable,
    title = "Composable Sparse Fine-Tuning for Cross-Lingual Transfer",
    author = "Ansell, Alan  and
      Ponti, Edoardo  and
      Korhonen, Anna  and
      Vuli{\'c}, Ivan",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.125/",
    doi = "10.18653/v1/2022.acl-long.125",
    pages = "1778--1796"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/