MMTEB: Massive Multilingual Text Embedding Benchmark
Kenneth C. EnevoldsenIsaac ChungImene KerbouaMrton KardosAshwin MathurDavid StapJay GalaWissam SibliniDominik KrzeminskiGenta Indra Winata
Presents the Massive Multilingual Text Embedding Benchmark (MMTEB), a suite of over 500 evaluation tasks across 250+ languages that incorporates efficient downsampling methods to drastically cut compute costs while revealing that compact 560M-parameter models can outperform multi-billion-parameter language models.
Text embeddings form the foundation for critical language applications, including semantic search, document classification, and retrieval-augmented generation. Despite their broad adoption, existing embedding benchmarks historically suffered from major limitations: they were constrained to a narrow set of tasks, focused heavily on English, and imposed prohibitive computational burdens. These resource demands disproportionately excluded mid- and low-resource language communities from evaluating and developing effective language technologies.
The article introduces the Massive Multilingual Text Embedding Benchmark (MMTEB) to provide a truly comprehensive, accessible, and high-quality evaluation framework. MMTEB establishes the largest multilingual benchmark collection to date, expanding prior efforts to cover over 500 tasks across 10 distinct task categories and more than 250 languages.
To construct the benchmark, an open, community-driven collaboration engaged native speakers and researchers globally. The framework integrates advanced challenges, such as instruction-following, long-document retrieval, code search, and multi-label classification. To resolve the computational bottleneck, the authors designed innovative optimization strategies, including hard-negative mining for retrieval, bootstrapping that reuses encoded text for clustering, and feature-selection techniques to eliminate redundant, highly correlated tasks.
The benchmark demonstrates several key findings:
- Instruction-tuned models drastically outperform standard architectures across most tasks, notably in bitext mining and clustering.
- Model size does not guarantee multilingual superiority: the 560-million-parameter multilingual-e5-large-instruct model consistently outperforms 7-billion-parameter models (such as Mistral-based GritLM-7B) on multilingual and low-resource language benchmarks, driven by more diverse cross-lingual pre-training.
- Large language models excel primarily on English and high-resource languages, but their relative advantage deteriorates sharply as the speaker count of a language decreases.
- Optimized benchmark subsets drastically reduce compute costs—cutting the zero-shot English evaluation down to 2% of the original document volume and reducing evaluation time to just 3.1 hours for a 7-billion-parameter model on a single graphics processing unit—while preserving model ranking accuracy (Spearman correlation of 0.90 to 0.97).
These findings carry significant strategic implications for engineering, cost, and language policy. Organizations deploying retrieval and search applications across global markets do not necessarily require multi-billion-parameter models; smaller, instruction-tuned multilingual models provide higher accuracy at a fraction of the serving and infrastructure costs. Furthermore, the dramatic compute reductions lower the barrier to entry, enabling teams with limited hardware budgets to reliably validate models in local and underserved languages.
Decision-makers and developers should adopt instruction-tuned, multilingual-first embeddings when building internationalized systems and utilize MMTEB's compact benchmark subsets for rapid, low-cost testing pipelines. Prior to deployment, teams should evaluate models using targeted regional benchmarks (such as MTEB Indic or Europe) rather than relying solely on high-resource English metrics.
The authors note limitations regarding potential English bias arising from human-translated data and an evaluation dataset distribution that remains skewed toward high-resource languages. Future efforts should expand authentic, natively generated datasets in underserved languages.
- Paper: Text Embeddings by Weakly-Supervised Contrastive Pre-training, Liang Wang et al. (2022). Introduces the foundational E5 text embedding architecture and the initial MTEB evaluation paradigm upon which MMTEB directly builds and scales.
- Paper: M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation, Jianlv Chen et al. (2024). Pioneers multi-lingual, multi-granularity retrieval embedding architectures that MMTEB systematically benchmarks across 250+ languages.
- Paper: Language-agnostic BERT Sentence Embedding, Fangxiaoyu Feng et al. (2020). Establishes large-scale dual-encoder pretraining across 100+ languages, providing the core multilingual embedding principles tested extensively in MMTEB.
- Paper: Making Monolingual Sentence Embeddings Multilingual Using Knowledge Distillation, Nils Reimers et al. (2020). Introduces multilingual knowledge distillation for sentence embeddings, establishing key baselines and evaluation concepts used in massively multilingual benchmarks.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). Provides the foundational cross-lingual NLI benchmark that established standardized cross-lingual evaluation paradigms adopted by MMTEB.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). Demonstrates the methods and challenges of scaling multilingual corpora and evaluation across 500+ diverse languages.
- Paper: SimCSE: Simple Contrastive Learning of Sentence Embeddings, Tianyu Gao et al. (2021). Presents foundational contrastive representation learning techniques for sentence embeddings that underpin modern embedders assessed by MMTEB.
- Paper: Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models, Yanzhao Zhang et al. (2025). Evaluates the next-generation Qwen3 embedding foundation models directly against the comprehensive multilingual benchmarks defined by MTEB and MMTEB.
- Paper: NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models, Chankyu Lee et al. (2025). Applies decoder-only LLM adaptation and contrastive instruction-tuning to achieve top-tier performance on large-scale MTEB evaluations.
- Paper: Making Text Embedders Few-Shot Learners, Chaofan Li et al. (2025). Extends dense text embedders to in-context few-shot learners and evaluates their performance across diverse task suites like MTEB.
- Paper: Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini, Madhuri Shanbhogue et al. (2026). Expands multilingual embedding evaluation beyond text into unified multimodal domains while validating against multilingual MTEB task suites.
- Paper: On the Theoretical Limitations of Embedding-Based Retrieval, Orion Weller et al. (2026). Analyzes the theoretical and practical retrieval capacity limitations of single-vector embedding models assessed in massive retrieval benchmarks like MMTEB.
