Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization

Negar ForoutanClara MeisterDebjit PaulJoel NiklausSina AhmadiAntoine BosselutRico Sennrich

article2025ACL20 citationsSAC Highlight (Area Chair's Award) at ACL 2026

Introduces Parity-aware Byte Pair Encoding, an algorithm that applies a fair-max merge rule to reduce cross-lingual tokenization disparity by up to 89% without compromising global compression rates or downstream language model performance.

Listen

Standard natural language processing pipelines rely on subword tokenization to convert raw text into model-readable units. However, predominant tokenization methods favor high-resource languages that dominate training data because they build vocabularies by maximizing overall corpus-level frequency. As a result, lower-resource and non-dominant languages suffer from poorer compression, meaning they are segmented into longer sequences of smaller fragments. This imbalance creates qualitative issues, such as degraded linguistic representation, alongside financial and technical burdens. Because commercial language model services typically bill and scale inference latency based on token counts, users of underrepresented languages effectively pay a disproportionate computational and financial tax.

The article introduces and evaluates Parity-Aware Byte-Pair Encoding, a modified tokenization approach designed to achieve cross-lingual equity in compression rates. The primary objective is to demonstrate that introducing a fair-selection rule during vocabulary learning reduces tokenization disparities across languages without sacrificing overall compression efficiency or downstream language model performance.

To address this challenge, the authors adapted the standard Byte-Pair Encoding framework. Instead of greedily selecting subword merges based solely on global dataset frequency, the fair-max approach identifies the language experiencing the worst compression rate at each merge step and selects the next merge based on frequency statistics from that specific language. The evaluation was conducted across datasets covering up to 60 diverse languages and multiple vocabulary sizes, utilizing both balanced and realistic unbalanced corpus distributions. Intrinsic properties—including compression rates, vocabulary utilization, and cross-language inequality measured by the Gini coefficient—were evaluated on parallel benchmarks. In addition, the researchers pretrained 3-billion-parameter language models across 100 billion tokens to conduct extrinsic zero-shot evaluations across 12 multilingual downstream tasks in 22 languages.

The investigation produced several key findings. First, the fair-max approach reduced cross-lingual tokenization inequality by up to 89% relative to standard Byte-Pair Encoding as measured by the Gini coefficient, producing much more uniform compression across languages. Second, this substantial fairness improvement incurred negligible loss in global compression rate and maintained vocabulary diversity and morphological alignment. Third, extrinsic evaluations demonstrated no systematic degradation in downstream model accuracy, with model performance remaining within statistical error bounds of the baseline across all tested languages. Finally, language models trained with fair-max tokenizers exhibited a more uniform distribution of language modeling perplexity across diverse languages, mitigating the extreme performance degradation seen on tail languages under the classical algorithm.

These findings indicate that tokenization disparities can be substantially mitigated without incurring operational drawbacks or performance penalties. For organizations deploying multilingual systems, adopting equitable tokenization directly lowers operational latency and service costs for underrepresented languages. The methodology acts as a direct drop-in replacement for standard tokenization pipelines: it adds minimal computational overhead during vocabulary construction and requires zero changes to the underlying model architecture or inference runtime. Pragmatic variants, such as hybrid scheduling and ratio-based targets, also allow practitioners to flexibly balance absolute compression against fairness according to domain-specific needs.

Organizations training multilingual language models should consider adopting parity-aware tokenization to eliminate structural cost and performance biases. In practice, implementation requires only a small parallel development set—as few as 100 aligned examples per language—or pre-estimated compression ratios to guide merge decisions. Future technical efforts should extend this fair-max criterion to alternative tokenization algorithms such as UnigramLM and evaluate performance in ultra-large-scale models, multimodal environments, and specialized domains such as software code.

Readers should note a few operational limitations when assessing these results. Standard implementations rely on parallel reference corpora, meaning performance may degrade if domain mismatch occurs between the reference data and pretraining distributions. Additionally, tokenization gains remain bounded by initial text segmentation rules, such as whitespace boundaries, which can limit compression potential in specific scripts. Nevertheless, empirical confidence in the reported findings is high across standard multilingual text applications, supported by consistent intrinsic and extrinsic validation across multiple languages, corpus sizes, and downstream benchmarks.

No sufficiently relevant recommendations were found.

Cover for Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization

Abstract

Tokenization is the first -- and often least scrutinized -- step of most NLP pipelines. Standard algorithms for learning tokenizers rely on frequency-based objectives, which favor languages dominant in the training data and consequently leave lower-resource languages with tokenizations that are disproportionately longer, morphologically implausible, or even riddled with <UNK><UNK> placeholders. This phenomenon ultimately amplifies computational and financial inequalities between users from different language backgrounds. To remedy this, we introduce Parity-aware Byte Pair Encoding (BPE), a variant of the widely-used BPE algorithm. At every merge step, Parity-aware BPE applies a fair-max rule that maximizes the compression gain of the currently worst-compressed language, trading a small amount of global compression for cross-lingual parity. We find empirically that Parity-aware BPE reduces tokenization inequality -- operationalized by the Gini coefficient of per-language token costs -- by up to 89% relative to Classical BPE. This comes with negligible impact on global compression rate and no evidence of systematic degradation in downstream LM performance.

Table of Contents

  • 1 Introduction
  • 2 Text Tokenization
  • 2.1 Byte-level Tokenizers
  • 2.2 Text Compression
  • 2.3 Tokenization Fairness
  • 3 Byte Pair Encoding
  • 4 Parity-aware Byte Pair Encoding
  • 4.1 Greedy Fair-Max Objective
  • 4.2 Algorithm
  • 4.3 Algorithmic Variants
  • 5 Experimental Setup
  • 5.1 Tokenizer Training
  • 5.2 Evaluation
  • 5.2.1 Intrinsic Metrics
  • 5.2.2 Extrinsic Metrics.
  • 6 Results and Analysis
  • 6.1 Intrinsic Evaluation
  • 6.2 Extrinsic Evaluation
  • 7 Related Work
  • 8 Discussion and Conclusion
  • References
  • A Pseudocode
  • B Intrinsic Tokenizer Evaluation Metrics
  • B.1 Vocabulary Usage
  • B.2 Information-theoretic Metrics
  • B.3 Morphological and Multilingual Fairness Metrics
  • C Hyperparameter Selection
  • C.1 Hybrid Parity-aware BPE
  • C.2 Moving-Window Balancing
  • C.3 Parallel Development Set Size
  • D Additional Results and Ablation Studies
  • E Balancing the tokenizer training set
  • F Language Model Training
  • F.1 Model Architecture
  • F.2 Training Hyperparameters
  • F.3 Hardware Setup
  • F.4 Sampling Methods
  • G Downstream Benchmark Evaluation
  • G.1 Benchmarks
  • G.2 Score Aggregations

Knowls

  1. Knowl 1 — Fair-max objective for multilingual BPE

    equation

    For a language set L\mathcal{L}, let DℓD_\ell be the byte-strings for language ℓ\ell, and let τm\tau_m be the tokenizer defined by an ordered BPE merge list mm. Define language compression rate as CR(ℓ;τm)=∑b∈Dℓ∣b∣u∑b∈Dℓ∣τm(b)∣CR(\ell;\tau_m)=\frac{\sum_{b\in D_\ell}|b|_u}{\sum_{b\in D_\ell}|\tau_m(b)|}, where ∣b∣u|b|_u counts units of text (for example, aligned lines or sentences) and ∣τm(b)∣|\tau_m(b)| counts output tokens. Parity-aware BPE targets a merge list of fixed length KK that maximizes the compression rate of the worst-compressed language:

    m⋆=arg⁡max⁡m:∣m∣=Kmin⁡ℓ∈LCR(ℓ;τm).m^\star=\arg\max_{m:|m|=K}\min_{\ell\in\mathcal{L}}CR(\ell;\tau_m).

    This is a max–min fairness target, not a claim that the greedy learning procedure finds the global optimum. It prioritizes improving the lowest language-level compression rate rather than maximizing corpus-wide compression.

  2. Knowl 2 — Greedy Parity-aware BPE learning procedure

    algorithm

    The algorithm changes BPE merge learning, not the resulting tokenizer’s inference procedure. It requires language-labeled training data for counting pair frequencies and recommends a separate development set with aligned content across languages for comparing compression rates; the training data itself need not be parallel. At each merge, it finds the currently lowest-compression language on the development set, counts adjacent token pairs only in that language’s training data, selects its most frequent pair, and applies the resulting merge to every language. It starts from the 256 singleton-byte vocabulary and stops after the requested number KK of merges.

    Input: Language-labeled training corpora {Dℓ}\{D_\ell\}, comparable development corpora {Dℓdev}\{D^{dev}_\ell\}, merge count KK
    Output: Vocabulary VKV_K and ordered merge list mKm_K
    Initialize V0V_0 to the singleton-byte vocabulary and m0m_0 to the empty list
    For kk from 1 to KK:
        For each language ℓ\ell, compute its compression rate on DℓdevD^{dev}_\ell using the current merges
        Set ℓ⋆\ell^\star to the language with the lowest development compression rate
        Count each adjacent token-pair type in the currently tokenized training corpus Dℓ⋆D_{\ell^\star}
        Select the pair (v⋆,v′⋆)(v^\star,v'^\star) with the highest count
        Add the concatenated token v⋆∘v′⋆v^\star\circ v'^\star to the vocabulary and append the pair to the merge list
        Apply this merge to all training and development corpora for every language
    Return VKV_K and mKm_K

    The merge selected from the focus language is applied across all languages, distinguishing this procedure from simply combining separate monolingual merge lists. The paper reports the same asymptotic complexity as Classical BPE and an additional O(∣L∣)O(|\mathcal{L}|) per-merge overhead for recomputing language compression rates on the development data.

  3. Knowl 3 — Gini coefficient for tokenization cost inequality

    definition

    The paper measures cross-language tokenization fairness as inequality in per-language token costs. For each language, cost is the average number of tokens used to encode a comparable content unit, such as a line or sentence in a parallel corpus. Let c1≤c2≤⋯≤cnc_1\leq c_2\leq\cdots\leq c_n be these costs for nn languages, sorted in ascending order. The token-cost Gini coefficient is

    Gini⁡=1n(n+1−2∑i=1n(n+1−i)ci∑i=1nci).\operatorname{Gini}=\frac{1}{n}\left(n+1-2\frac{\sum_{i=1}^{n}(n+1-i)c_i}{\sum_{i=1}^{n}c_i}\right).

    A value of zero indicates equal costs across languages; values closer to one indicate greater inequality. Using aligned content units makes the comparison less dependent on language-specific differences such as average bytes per character.

  4. Knowl 4 — Intrinsic fairness gains with nearly unchanged global compression

    data/table

    For the unbalanced 30-language FineWeb2 setup, 128k-vocabulary tokenizers were evaluated on the corresponding parallel FLORES+ language set. The table reports global statistics except MorphScore precision and recall, which are macro-averaged over languages with available measurements. The base Parity-aware BPE has the lowest Gini coefficient, falling from 0.064 for Classical BPE to 0.007, an approximately 89% reduction. Global compression rates remain close across variants; the ratios variant has a smaller fairness gain and lower vocabulary utilization than the other parity-aware variants.

    Tokenizer CR ↑\uparrow TTR ↑\uparrow Vocab. utilization ↑\uparrow Fertility ↓\downarrow Rényi (α=2.5\alpha=2.5) ↑\uparrow Gini ↓\downarrow MorphScore precision ↑\uparrow MorphScore recall ↑\uparrow
    Classical BPE 0.0275 0.0777 67.0% 3.990 0.49 0.064 0.537 0.659
    PA-BPE 0.0276 0.0819 70.4% 4.080 0.49 0.007 0.539 0.671
    PA-BPE (window) 0.0278 0.0838 71.6% 4.042 0.48 0.009 0.546 0.677
    PA-BPE (hybrid) 0.0278 0.0817 69.8% 4.056 0.49 0.015 0.541 0.667
    PA-BPE (hybrid+window) 0.0279 0.0832 70.8% 4.029 0.48 0.019 0.546 0.670
    PA-BPE (ratios) 0.0270 0.0737 64.9% 4.398 0.49 0.040 0.532 0.669
  5. Knowl 5 — Downstream models show no systematic performance degradation

    empirical result

    The downstream comparison trained 3-billion-parameter decoder-only Transformer models using 128k tokenizers and 100 billion FineWeb2 training tokens. It compared Classical BPE with hybrid Parity-aware BPE and hybrid Parity-aware BPE with moving-window balancing, evaluating across 12 multilingual benchmarks and 22 languages. The hybrid Parity-aware tokenizer produced nominal accuracy gains in 14 languages and nominal declines in 6; individual differences were generally within standard errors. The authors therefore found no evidence of systematic downstream degradation from using the parity-aware tokenizers. Per-language validation perplexity, normalized by bytes rather than tokens, also showed a larger cross-language spread for Classical BPE, while parity-aware variants appeared to reduce its high-perplexity tail with comparable mean perplexity.

  6. Knowl 6 — Hybrid learning trades global compression for parity

    model/method

    Hybrid Parity-aware BPE first learns JJ merges using Classical BPE’s global pair-frequency criterion, then learns another KK merges using the Parity-aware criterion, which selects pairs based on the currently worst-compressed language. Developers can vary the two merge counts to control the balance between corpus-wide compression and cross-language parity. In the reported experiments, the two phases each received half of the total merges; the paper does not claim that this split is universally optimal.

  7. Knowl 7 — Moving-window balancing limits repeated focus on one language

    model/method

    The moving-window variant tracks the languages selected as worst-compressed over the most recent WW merge steps. A language is excluded from selection if it appears more than αW/∣L∣\alpha W/|\mathcal{L}| times in that window; the next merge is then chosen using the eligible language with the lowest current compression rate. This safeguard is intended for cases where compression gains for a repeatedly selected language stagnate, for example because the development data poorly matches the training data or because few useful merges remain. The reported default settings are W=100W=100 and α=2\alpha=2. The mechanism can relax strict parity when a language is excluded.

  8. Knowl 8 — Ratio-normalized selection removes the development-set requirement

    model/method

    When aligned development data is unavailable or inappropriate, developers can assign each language a positive target compression ratio rℓr_\ell and compute compression rates on the training data. At each merge, the focus language is selected by

    ℓ⋆=arg⁡min⁡ℓ∈LCR(ℓ;τm<k)rℓ,\ell^\star=\arg\min_{\ell\in\mathcal{L}}\frac{CR(\ell;\tau_{m_{<k}})}{r_\ell},

    where m<km_{<k} is the merge list learned so far. The rule prioritizes the language whose current compression rate is lowest relative to its target. Equal targets express an equal-compression goal on the measured data; targets proportional to average bytes per line in a reference parallel corpus approximate content-based normalization. The authors recommend the latter to account for cross-language differences in UTF-8 size.

  9. Knowl 9 — Fairness gains persist with small development sets and varied windows

    empirical result

    Ablations found little sensitivity to moving-window sizes of 50, 100, 150, and 200: compression rates were 0.0277, 0.0278, 0.0277, and 0.0278, respectively, and Gini values were 0.008, 0.009, 0.009, and 0.009. Smaller aligned development sets also retained much of the fairness improvement: the paper reports a 67% Gini reduction relative to the Classical BPE baseline with 100 examples, compared with 72% using the full 1,000-example development set. Global compression rates remained essentially unchanged across development-set sizes.

  10. Knowl 10 — Limits of parity-aware tokenization

    limitation

    The standard selection rule depends on estimating comparable per-language compression rates, for which the authors recommend aligned development data. If such data is unavailable, ratio targets may have to be estimated heuristically, weakening identification of the truly worst-compressed language. BPE also cannot shorten a string beyond its pre-tokenized length, so languages near that limit may yield diminishing compression gains; the moving-window variant responds by relaxing the parity constraint, allowing disparities to remain. The experiments cover up to 60 languages and 3-billion-parameter models, leaving larger-scale, code, and multimodal settings untested. Fairness is defined primarily through token counts, so other possible notions of linguistic or processing equity are not optimized.

Coverage note — Detailed per-language benchmark scores and ancillary tokenizer diagnostics are omitted because their aggregate patterns are represented in the reported results; no other major contributed component is deliberately omitted.

References

  1. 1.Khadige Abboud and Gokmen Oz. 2024. Towards equitable natural language understanding systems for dialectal cohorts: Debiasing training data. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC/COLING 2024, 20-25 May, 2024, Torino, Italy, pages 16487–16499. ELRA and ICCL.
  2. 2.Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David R. Mortensen, Noah A. Smith, and Yulia Tsvetkov. 2023. Do all languages cost the same? tokenization in the era of commercial language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 9904–9923. Association for Computational Linguistics.
  3. 3.Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Buschhoff, Charvi Jain, Alexander Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suarez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, Stefan Kesselheim, and Nicolas Flores-Herr. 2024a. Tokenizer choice for LLM training: Negligible or crucial? In Findings of the Association for Computational Linguistics: NAACL 2024, pages 3907–3924, Mexico City, Mexico. Association for Computational Linguistics.
  4. 4.Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Buschhoff, et al. 2024b. Tokenizer choice for llm training: Negligible or crucial? In Findings of the Association for Computational Linguistics: NAACL 2024, pages 3907–3924.
  5. 5.Alejandro Hernández-Cano Apertus Project, Alexander Hägele, Allen Hao Huang, Angelika Romanou, Antoni-Joan Solergibert, Barna Pasztor, Bettina Messmer, Dhia Garbaya, Eduard Frank Durech, Ido ˇ Hakimi, Juan García Giraldo, Mete Ismayilzada, Negar Foroutan, Skander Moalla, Tiancheng Chen, Vinko Sabolcec, Yixuan Xu, Michael Aerni, Badr ˇ AlKhamissi, Ines Altemir Marinas, Mohammad Hossein Amani, Matin Ansaripour, Ilia Badanin, Harold Benoit, Emanuela Boros, Nicholas Browning, Fabian Bösch, Maximilian Böther, Niklas Canova, Camille Challier, Clement Charmillot, Jonathan Coles, Jan Deriu, Arnout Devos, Lukas Drescher, Daniil Dzenhaliou, Maud Ehrmann, Dongyang Fan, Simin Fan, Silin Gao, Miguel Gila, María Grandury, Diba Hashemi, Alexander Hoyle, Jiaming Jiang, Mark Klein, Andrei Kucharavy, Anastasiia Kucherenko, Frederike Lübeck, Roman Machacek, Theofilos Manitaras, Andreas Marfurt, Kyle Matoba, Simon Matrenok, Henrique Mendoncça, Fawzi Roberto Mohamed, Syrielle Montariol, Luca Mouchel, Sven Najem-Meyer, Jingwei Ni, Gennaro Oliva, Matteo Pagliardini, Elia Palme, Andrei Panferov, Léo Paoletti, Marco Passerini, Ivan Pavlov, Auguste Poiroux, Kaustubh Ponkshe, Nathan Ranchin, Javi Rando, Mathieu Sauser, Jakhongir Saydaliev, Muhammad Ali Sayfiddinov, Marian Schneider, Stefano Schuppli, Marco Scialanga, Andrei Semenov, Kumar Shridhar, Raghav Singhal, Anna Sotnikova, Alexander Sternfeld, Ayush Kumar Tarun, Paul Teiletche, Jannis Vamvas, Xiaozhe Yao, Hao Zhao Alexander Ilic, Ana Klimovic, Andreas Krause, Caglar Gulcehre, David Rosenthal, Elliott Ash, Florian Tramèr, Joost VandeVondele, Livio Veraldi, Martin Rajman, Thomas Schulthess, Torsten Hoefler, Antoine Bosselut, Martin Jaggi, and Imanol Schlag. 2025. Apertus: Democratizing open and compliant llms for global language environments. arXiv e-prints, pages arXiv–2509.
  6. 6.Catherine Arnett, Tyler A. Chang, Stella Biderman, and Benjamin K. Bergen. 2025a. Explaining and mitigating crosslingual tokenizer inequities. In Advances in Neural Information Processing Systems 38 (NeurIPS 2025).
  7. 7.Catherine Arnett, Marisa Hudspeth, and Brendan O’Connor. 2025b. Evaluating Morphological Alignment of Tokenizers in 70 Languages. In Proceedings of the ICML 2025 Tokenization Workshop (TokShop).
  8. 8.Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Kelly Marchisio, Sebastian Ruder, et al. 2024. Aya 23: Open weight releases to further multilingual progress. CoRR, abs/2405.15032.
  9. 9.Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. The Belebele benchmark: a parallel reading comprehension dataset in 122 language variants. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 749–775, Bangkok, Thailand. Association for Computational Linguistics.
  10. 10.Kaj Bostrom and Greg Durrett. 2020. Byte pair encoding is suboptimal for language model pretraining. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4617–4624, Online. Association for Computational Linguistics.
  11. 11.Michael Chen, Mike D’Arcy, Alisa Liu, Jared Fernandez, and Doug Downey. 2019. CODAH: An adversarially-authored question answering dataset for common sense. In Proceedings of the 3rd Workshop on Evaluating Vector Space Representations for NLP, pages 63–69. Association for Computational Linguistics.
  12. 12.Hyung Won Chung, Dan Garrette, Kiat Chuan Tan, and Jason Riesa. 2020. Improving multilingual models with language-clustered vocabularies. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 4536–4546. Association for Computational Linguistics.
  13. 13.Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting. 2022. CANINE: Pre-training an efficient tokenization-free encoder for language representation. Trans. Assoc. Comput. Linguistics, 10:73–91.
  14. 14.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451. Association for Computational Linguistics.
  15. 15.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  16. 16.Gautier Dagan, Gabriel Synnaeve, and Baptiste Rozière. 2024. Getting the most out of your tokenizer for pre-training and domain adaptation. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org.
  17. 17.Kazuki Fujii, Taishi Nakamura, Mengsay Loem, Hiroki Iida, Masanari Ohi, Kakeru Hattori, Hirai Shota, Sakae Mizuki, Rio Yokota, and Naoaki Okazaki. 2024. Continual pre-training for cross-lingual LLM adaptation: Enhancing japanese language capabilities. In First Conference on Language Modeling.
  18. 18.Philip Gage. 1994. A new algorithm for data compression. C Users Journal, 12(2):23–38.
  19. 19.Omer Goldman, Avi Caciularu, Matan Eyal, Kris Cao, Idan Szpektor, and Reut Tsarfaty. 2024. Unpacking tokenization: Evaluating text compression and its correlation with model performance. In Findings of the Association for Computational Linguistics: ACL 2024, pages 2274–2286, Bangkok, Thailand. Association for Computational Linguistics.
  20. 20.Thamme Gowda and Jonathan May. 2020. Finding the optimal vocabulary size for neural machine translation. In Findings of the Association for Computational Linguistics: EMNLP, volume EMNLP 2020 of Findings of ACL, pages 3955–3964. Association for Computational Linguistics.
  21. 21.Nathan Habib, Clémentine Fourrier, Hynek Kydlícek, Thomas Wolf, and Lewis Tunstall. 2023. Lighteval: A lightweight framework for llm evaluation.
  22. 22.Alexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben allal, Leandro Von Werra, and Martin Jaggi. 2024. Scaling laws and compute-optimal training beyond fixed training durations. In The Thirty-eighth Annual Conference on Neural Information Processing Systems.
  23. 23.Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, and Preslav Nakov. 2020. EXAMS: A multi-subject high school examinations dataset for cross-lingual and multilingual question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5427–5444, Online. Association for Computational Linguistics.
  24. 24.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR).
  25. 25.Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2022. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, 10:50–72.
  26. 26.Taku Kudo. 2018. Subword regularization: Improving neural network translation models with multiple subword candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL, pages 66–75. Association for Computational Linguistics.
  27. 27.Viet Lai, Chien Nguyen, Nghia Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan Rossi, and Thien Nguyen. 2023. Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 318–327. Association for Computational Linguistics.
  28. 28.Pietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos, and Tiago Pimentel. 2025. Causal estimation of tokenisation bias. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 28325–28340, Vienna, Austria. Association for Computational Linguistics.
  29. 29.Tomasz Limisiewicz, Jiří Balhar, and David Marecek. 2023. Tokenization impacts multilingual language modeling: Assessing vocabulary allocation and overlap across languages. In Findings of the Association for Computational Linguistics, pages 5661–5681. Association for Computational Linguistics.
  30. 30.Tomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia, and Luke Zettlemoyer. 2024. MYTE: Morphology-driven byte encoding for better and fairer multilingual language modeling. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, pages 15059–15076. Association for Computational Linguistics.
  31. 31.Bill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, and Xiang Ren. 2021. Common sense beyond English: Evaluating and improving multilingual language models for commonsense reasoning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1274–1287, Online. Association for Computational Linguistics.
  32. 32.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022a. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics.
  33. 33.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona T. Diab, Veselin Stoyanov, and Xian Li. 2022b. Few-shot learning with multilingual generative language models. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 9019–9052. Association for Computational Linguistics.
  34. 34.Alisa Liu, Jonathan Hayase, Valentin Hofmann, Sewoong Oh, Noah A. Smith, and Yejin Choi. 2026. SuperBPE: Space travel for language models. In Tokenization Workshop.
  35. 35.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations.
  36. 36.Pedro Henrique Martins, João Alves, Patrick Fernandes, Nuno M Guerreiro, Ricardo Rei, Amin Farajian, Mateusz Klimaszewski, Duarte M Alves, José Pombal, Nicolas Boizard, et al. 2025. Eurollm-9b: Technical report. arXiv preprint arXiv:2506.04079.
  37. 37.Clara Meister. 2025. TokEval: A tokenizer analysis suite.
  38. 38.Nasrin Mostafazadeh, Michael Roth, Annie Louis, Nathanael Chambers, and James Allen. 2017. LSDSem 2017 shared task: The story cloze test. In Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential and Discourse-level Semantics, pages 46–51.
  39. 39.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023. Crosslingual generalization through multitask finetuning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15991–16111, Toronto, Canada. Association for Computational Linguistics.
  40. 40.NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2024. Scaling neural machine translation to 200 languages. Nature, 630(8018):841–846.
  41. 41.Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason E. Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, and Srini Iyer. 2025. Byte latent transformer: Patches scale better than tokens. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9238–9258. ACL.
  42. 42.Guilherme Penedo, Hynek Kydlícek, Vinko Sabolcec, Bettina Messmer, Negar Foroutan, Amir Hossein Kargaran, Colin Raffel, Martin Jaggi, Leandro Von Werra, and Thomas Wolf. 2025. Fineweb2: One pipeline to scale them all — adapting pre-training data processing to every language. In Second Conference on Language Modeling.
  43. 43.Aleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, and Adel Bibi. 2023. Language model tokenizers introduce unfairness between languages. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023.
  44. 44.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020a. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(1).
  45. 45.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020b. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67.
  46. 46.Angelika Romanou, Negar Foroutan, Anna Sotnikova, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Maheshwary, Micol Altomare, Zeming Chen, Mohamed A. Haggag, Snegha A, Alfonso Amayuelas, Azril Hafizi Amirudin, Danylo Boiko, Michael Chang, Jenny Chim, Gal Cohen, Aditya Kumar Dalmia, Abraham Diress, Sharad Duwal, Daniil Dzenhaliou, Daniel Fernando Erazo Florez, Fabian Farestam, Joseph Marvin Imperial, Shayekh Bin Islam, Perttu Isotalo, Maral Jabbarishiviari, Börje F. Karlsson, Eldar Khalilov, Christopher Klamm, Fajri Koto, Dominik Krzeminski, Gabriel Adriano de Melo, Syrielle Montariol, Yiyang Nan, Joel Niklaus, Jekaterina Novikova, Johan Samir Obando Ceron, Debjit Paul, Esther Ploeger, Jebish Purbey, Swati Rajwal, Selvan Sunitha Ravi, Sara Rydell, Roshan Santhosh, Drishti Sharma, Marjana Prifti Skenduli, Arshia Soltani Moakhar, Bardia Soltani Moakhar, Ayush Kumar Tarun, Azmine Toushik Wasi, Thenuka Ovin Weerasinghe, Serhan Yilmaz, Mike Zhang, Imanol Schlag, Marzieh Fadaee, Sara Hooker, and Antoine Bosselut. 2025. INCLUDE: Evaluating multilingual language understanding with regional knowledge. In The Thirteenth International Conference on Learning Representations.
  47. 47.Phillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder, and Iryna Gurevych. 2021. How good is your tokenizer? on the monolingual performance of multilingual language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, pages 3118–3135. ACL.
  48. 48.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. WinoGrande: an adversarial winograd schema challenge at scale. Commun. ACM, 64(9):99–106.
  49. 49.Craig W Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris Tanner. 2024. Tokenization is more than compression. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 678–702. Association for Computational Linguistics.
  50. 50.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics.
  51. 51.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1, pages 4149–4158. Association for Computational Linguistics.
  52. 52.Alexey Tikhonov and Max Ryabinin. 2021. It’s all in the heads: Using attention heads as a baseline for cross-lingual transfer in commonsense reasoning. In Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, volume ACL/IJCNLP 2021 of Findings of ACL, pages 3534–3546. Association for Computational Linguistics.
  53. 53.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. LLaMA: Open and efficient foundation language models. CoRR, abs/2302.13971.
  54. 54.Ahmet Üstün, Viraat Aryabumi, Zheng Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, et al. 2024. Aya model: An instruction finetuned open-access multilingual language model. In Proceedings of the 62nd annual meeting of the Association for Computational Linguistics (volume 1: Long papers), pages 15894–15939.
  55. 55.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems, pages 5998–6008.
  56. 56.Ada Wan. 2022. Fairness in representation for multilingual NLP: Insights from controlled experiments on conditional language modeling. In International Conference on Learning Representations.
  57. 57.Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022. ByT5: Towards a token-free future with pre-trained byte-to-byte models. Trans. Assoc. Comput. Linguistics, 10:291–306.
  58. 58.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021a. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  59. 59.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021b. mt5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: Human language technologies, pages 483–498.
  60. 60.Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019. PAWS-X: A cross-lingual adversarial dataset for paraphrase identification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3687–3692, Hong Kong, China. Association for Computational Linguistics.
  61. 61.Shiyue Zhang, Vishrav Chaudhary, Naman Goyal, James Cross, Guillaume Wenzek, Mohit Bansal, and Francisco Guzmán. 2022. How robust is neural machine translation to language imbalance in multilingual tokenizer training? In Proceedings of the 15th biennial conference of the Association for Machine Translation in the Americas (Volume 1: Research Track), AMTA 2022, Orlando, USA, September 12-16, 2022, pages 97–116. Association for Machine Translation in the Americas.
  62. 62.Wenxuan Zhang, Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. 2023. M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models. Advances in Neural Information Processing Systems, 36:5484–5505.
  63. 63.Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell. 2023a. Tokenization and the noiseless channel. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5184–5207, Toronto, Canada. Association for Computational Linguistics.
  64. 64.Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Tim Vieira, Mrinmaya Sachan, and Ryan Cotterell. 2023b. A formal perspective on byte-pair encoding. In Findings of the Association for Computational Linguistics: ACL 2023, pages 598–614, Toronto, Canada. Association for Computational Linguistics.

Citation

MLA
Foroutan, N., et al. “Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization”. arXiv, 2025, http://arxiv.org/abs/2508.04796v3.
APA
Foroutan, N., Meister, C., Paul, D., Niklaus, J., Ahmadi, S., Bosselut, A., & Sennrich, R. (2025). Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization. arXiv. http://arxiv.org/abs/2508.04796v3
Chicago
Foroutan, N., C. Meister, D. Paul, et al. 2025. “Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization”. arXiv. http://arxiv.org/abs/2508.04796v3.
Harvard
Foroutan, N. et al. (2025) “Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2508.04796v3.
Vancouver
1. Foroutan N, Meister C, Paul D, Niklaus J, Ahmadi S, Bosselut A, Sennrich R (2025) Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization. arXiv

BibTeX

@article{foroutan2025parity,
  title = {Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization},
  author = {Foroutan, Negar and Meister, Clara and Paul, Debjit and Niklaus, Joel and Ahmadi, Sina and Bosselut, Antoine and Sennrich, Rico},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2508.04796v3},
  eprint = {2508.04796}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/