MMTEB: Massive Multilingual Text Embedding Benchmark

Kenneth C. EnevoldsenIsaac ChungImene KerbouaMrton KardosAshwin MathurDavid StapJay GalaWissam SibliniDominik KrzeminskiGenta Indra Winata

article2025ICLR101 citations

Presents the Massive Multilingual Text Embedding Benchmark (MMTEB), a suite of over 500 evaluation tasks across 250+ languages that incorporates efficient downsampling methods to drastically cut compute costs while revealing that compact 560M-parameter models can outperform multi-billion-parameter language models.

Listen

Text embeddings form the foundation for critical language applications, including semantic search, document classification, and retrieval-augmented generation. Despite their broad adoption, existing embedding benchmarks historically suffered from major limitations: they were constrained to a narrow set of tasks, focused heavily on English, and imposed prohibitive computational burdens. These resource demands disproportionately excluded mid- and low-resource language communities from evaluating and developing effective language technologies.

The article introduces the Massive Multilingual Text Embedding Benchmark (MMTEB) to provide a truly comprehensive, accessible, and high-quality evaluation framework. MMTEB establishes the largest multilingual benchmark collection to date, expanding prior efforts to cover over 500 tasks across 10 distinct task categories and more than 250 languages.

To construct the benchmark, an open, community-driven collaboration engaged native speakers and researchers globally. The framework integrates advanced challenges, such as instruction-following, long-document retrieval, code search, and multi-label classification. To resolve the computational bottleneck, the authors designed innovative optimization strategies, including hard-negative mining for retrieval, bootstrapping that reuses encoded text for clustering, and feature-selection techniques to eliminate redundant, highly correlated tasks.

The benchmark demonstrates several key findings:

  1. Instruction-tuned models drastically outperform standard architectures across most tasks, notably in bitext mining and clustering.
  2. Model size does not guarantee multilingual superiority: the 560-million-parameter multilingual-e5-large-instruct model consistently outperforms 7-billion-parameter models (such as Mistral-based GritLM-7B) on multilingual and low-resource language benchmarks, driven by more diverse cross-lingual pre-training.
  3. Large language models excel primarily on English and high-resource languages, but their relative advantage deteriorates sharply as the speaker count of a language decreases.
  4. Optimized benchmark subsets drastically reduce compute costs—cutting the zero-shot English evaluation down to 2% of the original document volume and reducing evaluation time to just 3.1 hours for a 7-billion-parameter model on a single graphics processing unit—while preserving model ranking accuracy (Spearman correlation of 0.90 to 0.97).

These findings carry significant strategic implications for engineering, cost, and language policy. Organizations deploying retrieval and search applications across global markets do not necessarily require multi-billion-parameter models; smaller, instruction-tuned multilingual models provide higher accuracy at a fraction of the serving and infrastructure costs. Furthermore, the dramatic compute reductions lower the barrier to entry, enabling teams with limited hardware budgets to reliably validate models in local and underserved languages.

Decision-makers and developers should adopt instruction-tuned, multilingual-first embeddings when building internationalized systems and utilize MMTEB's compact benchmark subsets for rapid, low-cost testing pipelines. Prior to deployment, teams should evaluate models using targeted regional benchmarks (such as MTEB Indic or Europe) rather than relying solely on high-resource English metrics.

The authors note limitations regarding potential English bias arising from human-translated data and an evaluation dataset distribution that remains skewed toward high-resource languages. Future efforts should expand authentic, natively generated datasets in underserved languages.

Cover for MMTEB: Massive Multilingual Text Embedding Benchmark

Abstract

Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more comprehensive evaluation, we introduce the Massive Multilingual Text Embedding Benchmark (MMTEB) - a large-scale, community-driven expansion of MTEB, covering over 500 quality-controlled evaluation tasks across 250+ languages. MMTEB includes a diverse set of challenging, novel tasks such as instruction following, long-document retrieval, and code retrieval, representing the largest multilingual collection of evaluation tasks for embedding models to date. Using this collection, we develop several highly multilingual benchmarks, which we use to evaluate a representative set of models. We find that while large language models (LLMs) with billions of parameters can achieve state-of-the-art performance on certain language subsets and task categories, the best-performing publicly available model is multilingual-e5-large-instruct with only 560 million parameters. To facilitate accessibility and reduce computational cost, we introduce a novel downsampling method based on inter-task correlation, ensuring a diverse selection while preserving relative model rankings. Furthermore, we optimize tasks such as retrieval by sampling hard negatives, creating smaller but effective splits. These optimizations allow us to introduce benchmarks that drastically reduce computational demands. For instance, our newly introduced zero-shot English benchmark maintains a ranking order similar to the full-scale version but at a fraction of the computational cost.

Table of Contents

  • 1 Introduction
  • 2 MMTEB Construction
  • 2.1 Open science effort
  • 2.2 Ensuring task quality
  • 2.3 Accessibility and benchmark optimization
  • 2.3.1 Downsampling and caching embeddings
  • 2.3.2 Encouraging smaller dataset submissions
  • 2.3.3 Task Selection
  • 2.4 Benchmark construction
  • 3 Experimental Settings
  • 3.1 Models
  • 3.2 Evaluation Scores
  • 3.3 Multilingual performance
  • 4 Analysis and Discussion
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Contributions
  • B Overview and Construction of Tasks
  • B.1 Introduction to benchmark tasks
  • B.2 Task construction
  • B.3 Novel datasets
  • B.4 Task Metadata
  • B.4.1 Domains
  • C Benchmark Optimizations
  • C.1 Speeding Up Tasks
  • C.1.1 Clustering
  • C.1.2 Retrieval
  • C.2 Code Optimizations
  • D Task Overview
  • D.1 Tasks
  • D.2 Languages
  • D.3 Examples
  • E Full results
  • E.1 Performance per Number of Speakers
  • F New Metrics
  • F.1 Abstention for retrieval and reranking tasks
  • G Models
  • H Benchmark Construction and Overview
  • H.1 Benchmark creation
  • H.2 Benchmark task overview
  • H.3 Performance on MTEB(eng, v2)
  • H.4 Performance on MTEB(Code)

Knowls

  1. Knowl 1 — Task Selection via Backward Predictability Elimination

    algorithm

    To build concise, diverse, and computationally manageable benchmark subsets from a large candidate pool of tasks T\mathcal{T}, benchmark creators can formulate task selection as a feature elimination problem. Given an initial set of evaluation tasks T\mathcal{T} and a set of representative embedding models M\mathcal{M}, the objective is to eliminate redundant tasks whose scores are linearly predictable from other tasks while preserving language diversity, task category coverage, and ranking stability.

    Input: Candidate task set T\mathcal{T}, evaluated models M\mathcal{M}, performance matrix S={st,m∣t∈T,m∈M}S = \{s_{t, m} \mid t \in \mathcal{T}, m \in \mathcal{M}\}, correlation threshold θ\theta, maximum error standard deviation tolerance ϵ\epsilon
    Output: Selected task subset Tsel\mathcal{T}_{\text{sel}}
    Tsel←T\mathcal{T}_{\text{sel}} \leftarrow \mathcal{T}
    repeat
        for each task t∈Tselt \in \mathcal{T}_{\text{sel}} do
            for each model mi∈Mm_i \in \mathcal{M} do
                Train linear regression estimator ft,mif_{t, m_i} on feature vectors xm=(sj,m)j∈Tsel∖{t}\mathbf{x}_{m} = (s_{j, m})_{j \in \mathcal{T}_{\text{sel}} \setminus \{t\}} for all m∈M∖{mi}m \in \mathcal{M} \setminus \{m_i\} to predict target st,ms_{t, m}
                Predict score s^t,mi=ft,mi(xmi)\hat{s}_{t, m_i} = f_{t, m_i}(\mathbf{x}_{m_i})
            Compute Spearman correlation rt=Spearman(st,s^t)r_t = \text{Spearman}(\mathbf{s}_t, \mathbf{\hat{s}}_t) across all models m∈Mm \in \mathcal{M}
            Compute mean squared error normalized by score standard deviation: NMSEt=MSE(st,s^t)σ(st)\text{NMSE}_t = \frac{\text{MSE}(\mathbf{s}_t, \mathbf{\hat{s}}_t)}{\sigma(\mathbf{s}_t)}
        
        Identify candidate task t∗=arg⁡max⁡t∈Tselrtt^* = \arg\max_{t \in \mathcal{T}_{\text{sel}}} r_t
        
        if rt∗<θr_{t^*} < \theta then
            break
        
        if removing t∗t^* would eliminate a language from its task category in Tsel\mathcal{T}_{\text{sel}} then
            mark t∗t^* as non-removable and continue
        if NMSEt∗>ϵ\text{NMSE}_{t^*} > \epsilon then
            mark t∗t^* as non-removable and continue
        
        Tsel←Tsel∖{t∗}\mathcal{T}_{\text{sel}} \leftarrow \mathcal{T}_{\text{sel}} \setminus \{t^*\}
    until no removable tasks remain
    return Tsel\mathcal{T}_{\text{sel}}

    In benchmark configurations, θ=0.8\theta = 0.8 and ϵ=0.5\epsilon = 0.5 are applied for multilingual benchmarks (such as MTEB(Multilingual) and MTEB(Europe)), while θ=0.9\theta = 0.9 and ϵ=0.2\epsilon = 0.2 are used for zero-shot benchmarks (such as MTEB(eng, v2)).

  2. Knowl 2 — Multilingual Embedding Model Evaluation Across Multilingual, European, and Indic Benchmarks

    data/table

    The evaluation of twelve representative text embedding models across three multilingual benchmark suites—MTEB(Multilingual) (132 datasets), MTEB(Europe) (74 datasets), and MTEB(Indic) (23 datasets)—evaluates model rankings and task performance across eight categories: Bitext Mining (Btxt), Pair Classification (Pr Clf), Classification (Clf), Semantic Textual Similarity (STS), Retrieval (Rtrvl), Multilabel Classification (M. Clf), Clustering (Clust), and Reranking (Rrnk). Rankings are aggregated using the Borda count method (where higher Borda counts indicate higher overall consensus ranking across tasks).

    Model Borda Count All Avg Cat. Avg Btxt Pr Clf Clf STS Rtrvl M. Clf Clust Rrnk
    MTEB(Multilingual) (132 tasks)
    multilingual-e5-large-instruct 1375 63.2 62.1 80.1 80.9 64.9 76.8 57.1 22.9 51.5 62.6
    GritLM-7B 1258 60.9 60.1 70.5 79.9 61.8 73.3 58.3 22.8 50.5 63.8
    e5-mistral-7b-instruct 1233 60.3 59.9 70.6 81.1 60.3 74.0 55.8 22.2 51.4 63.8
    multilingual-e5-large 1109 58.6 58.2 71.7 79.0 59.9 73.5 54.1 21.3 42.9 62.8
    multilingual-e5-base 944 57.0 56.5 69.4 77.2 58.2 71.4 52.7 20.2 42.7 60.2
    multilingual-mpnet-base 830 52.0 51.1 52.1 81.2 55.1 69.7 39.8 16.4 41.1 53.4
    multilingual-e5-small 784 55.5 55.2 67.5 76.3 56.5 70.4 49.3 19.1 41.7 60.4
    LaBSE 719 52.1 51.9 76.4 76.0 54.6 65.3 33.2 20.1 39.2 50.2
    multilingual-MiniLM-L12 603 48.8 48.0 44.6 79.0 51.7 66.6 36.6 14.9 39.3 51.0
    all-mpnet-base 526 42.5 41.1 21.2 70.9 47.0 57.6 32.8 16.3 40.8 42.2
    all-MiniLM-L12 490 42.2 40.9 22.9 71.7 46.8 57.2 32.5 14.6 36.8 44.3
    all-MiniLM-L6 418 41.4 39.9 20.1 71.2 46.2 56.1 32.5 15.1 38.0 40.3
    MTEB(Europe) (74 tasks)
    GritLM-7B 757 63.0 62.7 90.4 89.9 64.7 76.1 57.1 17.6 45.3 60.3
    multilingual-e5-large-instruct 732 62.2 62.3 90.4 90.0 63.2 77.4 54.8 17.3 46.9 58.4
    e5-mistral-7b-instruct 725 61.7 61.9 89.6 91.2 62.9 76.5 53.6 15.5 46.5 59.8
    multilingual-e5-large 586 58.5 58.7 84.5 88.8 60.4 75.8 50.8 15.0 38.2 55.9
    multilingual-e5-base 499 57.2 57.5 84.1 87.4 57.9 73.7 50.2 14.9 38.2 53.9
    multilingual-mpnet-base 463 54.4 54.7 79.5 90.7 56.6 74.3 41.2 6.9 35.8 52.3
    multilingual-e5-small 399 55.0 55.7 80.9 86.4 56.1 71.6 46.1 14.0 36.5 54.1
    LaBSE 358 51.8 53.5 88.8 85.2 55.1 65.7 34.4 16.3 34.3 48.7
    multilingual-MiniLM-L12 328 51.7 52.4 77.0 88.9 52.7 72.5 37.6 5.7 34.4 50.2
    all-mpnet-base 310 44.7 44.7 29.8 80.5 49.2 63.9 37.3 10.9 36.2 49.6
    all-MiniLM-L12 292 44.4 44.1 32.1 81.5 49.2 64.2 36.2 7.6 32.5 49.2
    all-MiniLM-L6 237 43.4 43.2 27.2 80.2 47.8 62.7 37.3 8.8 33.6 47.7
    MTEB(Indic) (23 tasks)
    multilingual-e5-large-instruct 209 70.2 71.6 80.4 76.3 67.0 53.7 84.9 – 51.7 87.5
    multilingual-e5-large 188 66.4 65.1 77.7 75.1 64.7 43.9 82.6 – 25.6 86.0
    multilingual-e5-base 173 64.6 62.6 74.2 72.8 63.8 41.1 77.8 – 24.6 83.8
    multilingual-e5-small 164 64.7 63.2 73.7 73.8 63.8 40.8 76.8 – 29.1 84.4
    GritLM-7B 151 60.2 58.0 58.4 67.8 60.0 27.2 79.5 – 28.0 84.7
    e5-mistral-7b-instruct 144 60.0 58.4 59.1 73.0 59.6 23.0 77.3 – 32.7 84.4
    LaBSE 139 61.9 59.7 74.1 64.6 61.9 52.8 64.3 – 21.1 79.0
    multilingual-mpnet-base 137 58.5 55.2 44.2 82.0 61.9 34.1 57.9 – 32.1 74.3
    multilingual-MiniLM-L12 98 49.7 42.2 15.3 77.8 57.6 19.8 48.8 – 16.7 59.3
    all-mpnet-base 68 33.6 22.6 3.7 52.6 45.2 -2.5 12.9 – 4.0 42.6
    all-MiniLM-L12 49 33.1 23.2 3.5 55.0 43.9 -5.3 13.9 – 3.7 47.6
    all-MiniLM-L6 40 31.8 20.4 2.5 53.7 44.1 -6.3 6.2 – 3.1 39.2

    The 560M parameter model multilingual-e5-large-instruct achieves the highest overall Borda rank on MTEB(Multilingual) and MTEB(Indic), outperforming 7B parameter models (GritLM-7B and e5-mistral-7b-instruct). On European languages (MTEB(Europe)), GritLM-7B achieves the highest score (63.0 average vs. 62.2 for multilingual-e5-large-instruct). Models trained exclusively on English data (all-mpnet-base, all-MiniLM-L12, all-MiniLM-L6) suffer severe degradation on non-English benchmarks.

  3. Knowl 3 — Normalized Area Under the Metric-Abstention Curve (nAUC)

    equation

    To evaluate score calibration and the ability of retrieval and reranking models to selectively abstain from making predictions on low-confidence queries, performance is measured using the normalized Area Under the metric-abstention Curve (nAUCmodeln\text{AUC}_{\text{model}}):

    nAUCmodel=AUCmodel−AUC−AUC+−AUC−n\text{AUC}_{\text{model}} = \frac{\text{AUC}_{\text{model}} - \text{AUC}^-}{\text{AUC}^+ - \text{AUC}^-}

    where:

    • An evaluation instance is a tuple (q,d1,…,dk)(q, d_1, \dots, d_k) containing a query qq and kk candidate documents d1,…,dkd_1, \dots, d_k.
    • c(q,d1,…,dk)∈Rc(q, d_1, \dots, d_k) \in \mathbb{R} is a confidence function derived from model similarity scores (e.g., maximum similarity score, score standard deviation, or margin between highest and second-highest scores).
    • τ1<τ2<⋯<τn\tau_1 < \tau_2 < \dots < \tau_n is an increasing sequence of abstention thresholds regulating the proportion of queries on which the model abstains.
    • For each threshold τi\tau_i, a sub-dataset Si={(q,d1,…,dk)∈D∣c(q,d1,…,dk)≥τi}S_i = \{(q, d_1, \dots, d_k) \in \mathcal{D} \mid c(q, d_1, \dots, d_k) \ge \tau_i\} is constructed, and evaluation metric m(Si)m(S_i) (such as nDCG@10 or MAP@1000) is computed.
    • AUCmodel\text{AUC}_{\text{model}} is the area under the resulting metric-versus-abstention-rate curve across all τi\tau_i.
    • AUC−\text{AUC}^- is the effective lower bound representing the area under the curve if task performance remains completely flat (unimproved) as abstention rate increases.
    • AUC+\text{AUC}^+ is the theoretical upper bound achieved by an oracle system with ground-truth access that discards incorrect predictions first at every abstention rate.
  4. Knowl 4 — Hard-Negative Truncation and Query Subsampling for Retrieval Evaluation

    model/method

    To drastically reduce the computational expense of evaluating text embedding models on large retrieval corpora containing millions of documents, retrieval tasks are downsampled using hard-negative pooling:

    1. Query Subsampling: For datasets exceeding 1,000 evaluation queries, a random subset of 1,000 queries is sampled.
    2. Candidate Pooling: For each sampled query, the top 250 ranked candidate documents are retained using a pool of negative-mining models (BM25 for lexical hard negatives, e5-multilingual-large for multilingual dense retrieval, and e5-mistral-7b-instruct for instruction-tuned LLM retrieval).
    3. Corpus Merging: All mined top-250 candidate lists across all queries, along with ground-truth positive documents, are merged and deduplicated to form a compact corpus of at most 250,000 documents.

    Downsampling to 250 documents per query reduces large corpus sizes (e.g., reducing 5 million documents down to ≤250,000\le 250,000, which is 2%2\% of the original document volume and 6%6\% of character count), while maintaining stable model rankings and preventing artificial inflation of nDCG@10 retrieval scores compared to the unpruned collection.

  5. Knowl 5 — Bootstrapped Single-Pass Corpus Subsampling for Clustering Evaluation

    model/method

    Traditional MTEB clustering evaluations repeatedly sample and encode separate subsets of documents across multiple runs, leading to redundant embedding computation. An optimized bootstrapping procedure reduces this burden:

    1. Corpus Subsampling: A single 4% random subset of the corpus is sampled (stratified by target label categories) and embedded once by the model.
    2. Bootstrap Resampling: From this pre-computed pool of embeddings, 10 distinct evaluation sets of size kk are sampled without replacement.
    3. Clustering and Evaluation: KK-means clustering is fitted independently on each of the 10 sampled sets, and clustering quality is computed against ground-truth labels using the V-measure.
    4. Aggregation: The final score is the mean V-measure across the 10 runs.

    Across 9 standard clustering tasks (Biorxiv P2P/S2S, Medrxiv P2P/S2S, Reddit P2P/S2S, StackExchange P2P/S2S, and TwentyNewsgroups), this method achieves an average evaluation speedup of 16.11×16.11\times while maintaining a mean Spearman rank correlation of 0.9634 relative to full evaluation rankings.

  6. Knowl 6 — Impact of Pre-Training Breadth Versus Model Scale Across Language Resource Levels

    empirical result

    Evaluating large language model-based embedders against smaller pre-trained multilingual encoders reveals that pre-training linguistic coverage dominates parameter scale on low- and mid-resource languages:

    • High-Resource / European Languages: 7B parameter models based on Mistral-7B (GritLM-7B and e5-mistral-7b-instruct) perform strongly, with GritLM-7B achieving the highest Borda count (757) and task average (63.0) on MTEB(Europe).
    • Low- and Mid-Resource Languages (<300M Speakers): The 560M parameter multilingual-e5-large-instruct (based on XLM-RoBERTa Large, pre-trained on 100 languages) significantly outperforms both 7B parameter models. On MTEB(Indic), multilingual-e5-large-instruct achieves an overall average of 70.2 compared to 60.2 for GritLM-7B and 60.0 for e5-mistral-7b-instruct.
    • Instruction-Tuning Effect: Fine-tuning with instruction prompts yields major gains regardless of architecture, notably on bitext mining and clustering, without requiring per-task manual prompt tuning.
  7. Knowl 7 — Zero-Shot English Benchmark Optimization and Performance (MTEB eng v2)

    data/table

    The MTEB(eng, v2) benchmark provides an accelerated, zero-shot evaluation of English text embeddings. It reduces the task count from 56 (in v1) to 41 tasks by applying task selection, while removing datasets commonly included in model training mixtures (such as MS MARCO and Natural Questions) to avoid fine-tuning leakage. It utilizes hard-negative downsampled retrieval and bootstrapped clustering splits, reducing document volume to 2%2\% of the original corpus.

    Model Borda Count All Avg Cat. Avg Pair Clf Clf STS Retrieval Clustering
    e5-mistral-7b-instruct 393 67.0 67.2 88.4 75.2 83.6 54.8 51.4
    GritLM-7B 384 66.4 66.7 87.3 77.0 82.5 53.2 50.8
    multilingual-e5-large-instruct 357 65.2 65.6 86.2 73.2 84.3 51.0 49.9
    multilingual-e5-large 270 62.1 62.4 84.7 72.8 80.6 49.0 42.8
    all-mpnet-base-v2 211 56.0 58.1 83.0 56.6 72.2 41.9 46.6
    multilingual-e5-base 211 60.2 60.9 83.6 70.0 79.1 46.1 42.2
    paraphrase-multilingual-mpnet-base-v2 188 57.3 58.8 81.7 68.6 79.8 34.1 43.5
    all-MiniLM-L12-v2 172 54.7 57.0 82.5 55.8 70.7 40.7 44.6
    all-MiniLM-L6-v2 149 54.4 56.7 82.4 55.4 70.4 39.8 44.9
    multilingual-e5-small 147 58.4 59.3 82.7 67.7 77.6 43.7 40.8
    paraphrase-multilingual-MiniLM-L12-v2 109 55.1 57.0 80.0 64.4 77.5 32.8 41.7
    LaBSE 49 48.6 51.7 78.9 66.8 70.2 16.8 36.1

    MTEB(eng, v2) maintains a Spearman correlation of 0.90 (p<0.0001p < 0.0001) and Pearson correlation of 0.96 (p<0.0001p < 0.0001) with MTEB(eng, v1) across evaluated models, while executing in 3.11 hours for a 7B LLM on a single H100 GPU.

  8. Knowl 8 — Code Information Retrieval Benchmark (MTEB Code) Model Performance

    data/table

    Evaluation of embedding models on MTEB(Code) / CoIR retrieval tasks across seven programming languages (C++, Go, Java, JavaScript, PHP, Python, Ruby) demonstrates the superiority of large generative models on code retrieval.

    Model Borda Rank (Count) All Avg C++ Go Java JavaScript PHP Python Ruby
    GritLM-7B 1 (88) 73.6 73.1 83.8 84.9 81.7 77.8 86.4 83.8
    e5-mistral-7b-instruct 2 (74) 69.2 68.3 83.0 80.9 79.4 75.6 83.6 81.1
    multilingual-e5-large-instruct 3 (65) 65.0 56.4 74.7 74.7 71.7 71.6 79.1 74.9
    multilingual-e5-large 4 (63) 61.7 46.8 73.4 72.2 66.6 69.1 75.7 73.4
    multilingual-e5-base 5 (55) 57.5 48.9 73.2 71.0 66.1 67.8 75.2 72.7
    multilingual-e5-small 6 (53) 58.4 48.4 70.6 67.9 65.2 66.6 73.6 68.1
    all-mpnet-base-v2 7 (44) 56.4 46.3 67.4 62.2 63.1 61.7 69.0 65.7
    all-MiniLM-L6-v2 8 (34) 52.7 48.1 64.4 57.4 62.2 60.4 68.1 66.6
    all-MiniLM-L12-v2 9 (27) 50.2 46.8 68.1 57.3 63.6 62.7 68.7 67.8
    LaBSE 10 (11) 28.8 27.6 40.6 36.6 42.3 34.8 43.9 42.2

    GritLM-7B achieves the highest average score (73.6) and dominates across all seven languages, followed by e5-mistral-7b-instruct (69.2) and multilingual-e5-large-instruct (65.0).

  9. Knowl 9 — Synthetic Multilingual Wikipedia Retrieval and Reranking Construction

    model/method

    To construct multilingual retrieval and reranking datasets across underrepresented languages (WikipediaRetrievalMultilingual and WikipediaRerankingMultilingual), a synthetic generation pipeline uses Wikipedia text and instruction-prompted LLMs:

    1. Article Sampling: Wikipedia articles with at least 9 paragraphs are identified from the top 100,000 viewed articles per language; 1,500 articles are sampled.
    2. Context Window Selection: A contiguous window of 9 paragraphs is selected from each article. The middle (5th) paragraph is designated as the positive document context d+d^+, and the remaining 8 neighboring paragraphs are extracted as local hard negatives {d1−,…,d8−}\{d^-_1, \dots, d^-_8\}.
    3. Query Generation: GPT-4o generates a realistic search question in the target language based on d+d^+, constrained to be precise, coherent without reading the text, single-aspect, non-overly-specific, and fully answerable by d+d^+.
    4. Task Formatting:
      • Reranking: Formatted as 1 positive context d+d^+ and 8 hard negative contexts {d1−,…,d8−}\{d^-_1, \dots, d^-_8\} per query.
      • Retrieval: Formatted with 1 positive context d+d^+, 8 local hard negatives, and all other non-target paragraphs in the sampled corpus as global negatives.

    Evaluation on German synthetic retrieval versus human-annotated GermanQuAD yields a Spearman rank correlation of 0.930.93 (95%95\% CI: [0.69,1.00][0.69, 1.00]).

  10. Knowl 10 — Instruction Retrieval and Multilabel Classification Task Formulations in MMTEB

    definition

    MMTEB introduces specific evaluation protocols for task categories previously unsupported or unstandardized in text embedding benchmarks:

    • Instruction Retrieval: Each search query qq is accompanied by a comprehensive, query-specific instruction IqI_q detailing explicit inclusion and exclusion relevance criteria (e.g., distinguishing valid treaty signatories from trade discussions). Systems must embed the query in the context of IqI_q to rank candidate documents. The primary evaluation metric is Robustness@10.
    • Multi-Label Classification: Documents may belong simultaneously to multiple classes. Evaluation uses 10-fold downsampled experiments with 8 training instances per unique class label. Embeddings are classified using a KK-Nearest Neighbors classifier and evaluated on held-out test splits using Accuracy, Macro-F1F_1, and Label Ranking Average Precision (LRAP).

Coverage note — Omitted individual descriptions of pre-existing datasets integrated into MMTEB (e.g., SIB-200, Belebele, NTREX, Flores) and standard subroutines in existing libraries, focusing on MMTEB's architectural contributions, downsampling methods, benchmark construction pipelines, and core empirical results.

References

  1. 1.David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba O Alabi, Yanke Mao, Haonan Gao, and Annie En-Shiun Lee. Sib-200: A simple, inclusive, and big evaluation dataset for topic classification in 200+ languages and dialects. arXiv preprint arXiv:2309.07445, 2023a.
  2. 2.David Ifeoluwa Adelani, Marek Masiak, Israel Abebe Azime, Jesujoba Alabi, Atnafu Lambebo Tonja, Christine Mwase, Odunayo Ogundepo, Bonaventure F. P. Dossou, Akintunde Oladipo, Doreen Nixdorf, Chris Chinenye Emezue, sana al azzawi, Blessing Sibanda, Davis David, Lolwethu Ndolela, Jonathan Mukiibi, Tunde Ajayi, Tatiana Moteu, Brian Odhiambo, Abraham Owodunni, Nnaemeka Obiefuna, Muhidin Mohamed, Shamsuddeen Hassan Muhammad, Teshome Mulugeta Ababu, Saheed Abdullahi Salahudeen, Mesay Gemeda Yigezu, Tajuddeen Gwadabe, Idris Abdulmumin, Mahlet Taye, Oluwabusayo Awoyomi, Iyanuoluwa Shode, Tolulope Adelani, Habiba Abdulganiyu, Abdul-Hakeem Omotayo, Adetola Adeeko, Abeeb Afolabi, Anuoluwapo Aremu, Olanrewaju Samuel, Clemencia Siro, Wangari Kimotho, Onyekachi Ogbu, Chinedu Mbonu, Chiamaka Chukwuneke, Samuel Fanijo, Jessica Ojo, Oyinkansola Awosan, Tadesse Kebede, Toadoum Sari Sakayo, Pamela Nyatsine, Freedmore Sidume, Oreen Yousuf, Mardiyyah Oduwole, Tshinu Tshinu, Ussen Kimanuka, Thina Diko, Siyanda Nxakama, Sinodos Nigusse, Abdulmejid Johar, Shafie Mohamed, Fuad Mire Hassan, Moges Ahmed Mehamed, Evrard Ngabire, Jules Jules, Ivan Ssenkungu, and Pontus Stenetorp. Masakhanews: News topic classification for african languages, 2023b.
  3. 3.Eneko Agirre, Mona Diab, Daniel Cer, and Aitor Gonzalez-Agirre. Semeval-2012 task 6: a pilot on semantic textual similarity. In Proceedings of the First Joint Conference on Lexical and Computational Semantics - Volume 1: Proceedings of the Main Conference and the Shared Task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation, SemEval ’12, pp. 385–393, USA, 2012. Association for Computational Linguistics.
  4. 4.Eneko Agirre, Daniel Matthew Cer, Mona T. Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. *sem 2013 shared task: Semantic textual similarity. In International Workshop on Semantic Evaluation, 2013. URL https://api.semanticscholar.org/CorpusID:10241043.
  5. 5.Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Inigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, et al. Semeval-2015 task 2: Semantic textual similarity, english, spanish and pilot on interpretability. In Proceedings of the 9th international workshop on semantic evaluation (SemEval 2015), pp. 252–263, 2015.
  6. 6.Gustavo Aguilar, Sudipta Kar, and Thamar Solorio. Lince: A centralized benchmark for linguistic code-switching evaluation. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pp. 1803–1813, 2020.
  7. 7.Orevaoghene Ahia, Julia Kreutzer, and Sara Hooker. The low-resource double bind: An empirical study of pruning for low-resource machine translation. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Findings of the Association for Computational Linguistics: EMNLP 2021, pp. 3316–3333, Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-emnlp.282. URL https://aclanthology.org/2021.findings-emnlp.282.
  8. 8.Vesa Akerman, David Baines, Damien Daspit, Ulf Hermjakob, Taeho Jang, Colin Leong, Michael Martin, Joel Mathew, Jonathan Robie, and Marcus Schwarting. The ebible corpus: Data and model benchmarks for bible translation for low-resource languages. arXiv preprint arXiv:2304.09919, 2023.
  9. 9.Dimosthenis Antypas, Asahi Ushio, Jose Camacho-Collados, Leonardo Neves, Vitor Silva, and Francesco Barbieri. Twitter Topic Classification. In Proceedings of the 29th International Conference on Computational Linguistics, Gyeongju, Republic of Korea, oct 2022. International Committee on Computational Linguistics.
  10. 10.Gaurav Arora. iNLTK: Natural language toolkit for indic languages. In Eunjeong L. Park, Masato Hagiwara, Dmitrijs Milajevs, Nelson F. Liu, Geeticka Chauhan, and Liling Tan (eds.), Proceedings of Second Workshop for NLP Open Source Software (NLP-OSS), pp. 66–71, Online, nov 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.nlposs-1.10. URL https://aclanthology.org/2020.nlposs-1.10.
  11. 11.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. On the cross-lingual transferability of monolingual representations. CoRR, abs/1910.11856, 2019.
  12. 12.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  13. 13.Soran Badawi, Arefeh Kazemi, and Vali Rezaie. Kurdisent: a corpus for kurdish sentiment analysis. Language Resources and Evaluation, pp. 1–20, 01 2024. doi: 10.1007/s10579-023-09716-6.
  14. 14.Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. The belebele benchmark: a parallel reading comprehension dataset in 122 language variants. arXiv preprint arXiv:2404.05961, 2023.
  15. 15.Anil Bandhakavi, Nirmalie Wiratunga, Deepak P, and Stewart Massie. Generating a word-emotion lexicon from #emotional tweets. In Johan Bos, Anette Frank, and Roberto Navigli (eds.), Proceedings of the Third Joint Conference on Lexical and Computational Semantics (*SEM 2014), pp. 12–21, Dublin, Ireland, aug 2014. Association for Computational Linguistics and Dublin City University. doi: 10.3115/v1/S14-1002. URL https://aclanthology.org/S14-1002.
  16. 16.Francesco Barbieri, Luis Espinosa Anke, and Jose Camacho-Collados. XLM-T: Multilingual language models in Twitter for sentiment analysis and beyond. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pp. 258–266, Marseille, France, jun 2022. European Language Resources Association. URL https://aclanthology.org/2022.lrec-1.27.
  17. 17.Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. Llm2vec: Large language models are secretly powerful text encoders. arXiv preprint arXiv:2404.05961, 2024.
  18. 18.Loubna Ben Allal, Niklas Muennighoff, Logesh Kumar Umapathi, Ben Lipkin, and Leandro von Werra. A framework for the evaluation of code generation models, 2022. URL https://github.com/bigcode-project/bigcode-evaluation-harness.
  19. 19.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen-tau Yih, and Yejin Choi. Abductive commonsense reasoning. In International Conference on Learning Representations, 2020.
  20. 20.Paheli Bhattacharya, Kripabandhu Ghosh, Saptarshi Ghosh, Arindam Pal, Parth Mehta, Arnab Bhattacharya, and Prasenjit Majumder. Aila 2019 precedent & statute retrieval task, oct 2020. URL https://doi.org/10.5281/zenodo.4063986.
  21. 21.Ergun Biçici. RTM-DCU: Predicting semantic similarity with referential translation machines. In Preslav Nakov, Torsten Zesch, Daniel Cer, and David Jurgens (eds.), Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval 2015), pp. 56–63, Denver, Colorado, jun 2015. Association for Computational Linguistics. doi: 10.18653/v1/S15-2010. URL https://aclanthology.org/S15-2010.
  22. 22.Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, et al. Lessons from the trenches on reproducible evaluation of language models. arXiv preprint arXiv:2405.14782, 2024.
  23. 23.Teven Le Scao BigScience Workshop, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic,´ Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, et al. Bloom: A 176b-parameter open-access multilingual language model, 2023.
  24. 24.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pp. 7432–7439, 2020.
  25. 25.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. Improving language models by retrieving from trillions of tokens, 2022.
  26. 26.Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. A full-text learning to rank dataset for medical information retrieval. 2016. URL http://www.cl.uni-heidelberg.de/~riezler/publications/papers/ECIR2016.pdf.
  27. 27.Chris Buckley, Darrin Dimmick, Ian Soboroff, and Ellen Voorhees. Bias and the limits of pooling for large collections. Information retrieval, 10:491–508, 2007.
  28. 28.Samuel Cahyawijaya, Holy Lovenia, Alham Fikri Aji, Genta Winata, Bryan Wilie, Fajri Koto, Rahmad Mahendra, Christian Wibisono, Ade Romadhony, Karissa Vincentio, et al. Nusacrowd: Open source initiative for indonesian nlp resources. In Findings of the Association for Computational Linguistics: ACL 2023, pp. 13745–13818, 2023a.
  29. 29.Samuel Cahyawijaya, Holy Lovenia, Fajri Koto, Dea Adhista, Emmanuel Dave, Sarah Oktavianti, Salsabil Akbar, Jhonson Lee, Nuur Shadieq, Tjeng Wawan Cenggoro, Hanung Linuwih, Bryan Wilie, Galih Muridan, Genta Winata, David Moeljadi, Alham Fikri Aji, Ayu Purwarianti, and Pascale Fung. Nusawrites: Constructing high-quality corpora for underrepresented and extremely low-resource languages. In Jong C. Park, Yuki Arase, Baotian Hu, Wei Lu, Derry Wijaya, Ayu Purwarianti, and Adila Alfa Krisnadhi (eds.), Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 921–945, Nusa Dua, Bali, nov 2023b. Association for Computational Linguistics. URL https://aclanthology.org/2023.ijcnlp-main.60.
  30. 30.Samuel Cahyawijaya, Holy Lovenia, Fajri Koto, Dea Adhista, Emmanuel Dave, Sarah Oktavianti, Salsabil Akbar, Jhonson Lee, Nuur Shadieq, Tjeng Wawan Cenggoro, et al. Nusawrites: Constructing high-quality corpora for underrepresented and extremely low-resource languages. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 921–945, 2023c.
  31. 31.Samuel Cahyawijaya, Holy Lovenia, Fajri Koto, Rifki Afina Putri, Emmanuel Dave, Jhonson Lee, Nuur Shadieq, Wawan Cenggoro, Salsabil Maulana Akbar, Muhammad Ihza Mahendra, et al. Cendol: Open instruction-tuned generative large language models for indonesian languages. arXiv preprint arXiv:2404.06138, 2024.
  32. 32.Iñigo Casanueva, Tadas Temcinas, Daniela Gerz, Matthew Henderson, and Ivan Vuli’c. Efficient intent detection with dual sentence encoders. In Tsung-Hsien Wen, Asli Celikyilmaz, Zhou Yu, Alexandros Papangelis, Mihail Eric, Anuj Kumar, Iñigo Casanueva, and Rushin Shah (eds.), Proceedings of the 2nd Workshop on Natural Language Processing for Conversational AI, pp. 38–45, Online, jul 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.nlp4convai-1.5. URL https://aclanthology.org/2020.nlp4convai-1.5.
  33. 33.Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In Steven Bethard, Marine Carpuat, Marianna Apidianaki, Saif M. Mohammad, Daniel Cer, and David Jurgens (eds.), Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pp. 1–14, Vancouver, Canada, aug 2017. Association for Computational Linguistics. doi: 10.18653/v1/S17-2001. URL https://aclanthology.org/S17-2001.
  34. 34.Ilias Chalkidis, Manos Fergadiotis, and Ion Androutsopoulos. Multieurlex – a multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2021. URL https://arxiv.org/abs/2109.00904.
  35. 35.Amit Kumar Chaudhary, Kurt Micallef, and Claudia Borg. Topic classification and headline generation for Maltese using a public news corpus. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation. Association for Computational Linguistics, may 2024.
  36. 36.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, et al. Evaluating large language models trained on code, 2021.
  37. 37.Xi Chen, Ali Zeynali, Chico Camargo, Fabian Fl"ock, Devin Gaffney, Przemyslaw Grabowicz, Scott Hale, David Jurgens, and Mattia Samory. SemEval-2022 task 8: Multilingual news article similarity. In Guy Emerson, Natalie Schluter, Gabriel Stanovsky, Ritesh Kumar, Alexis Palmer, Nathan Schneider, Siddharth Singh, and Shyam Ratan (eds.), Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022), pp. 1094–1106, Seattle, United States, jul 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.semeval-1.155. URL https://aclanthology.org/2022.semeval-1.155.
  38. 38.François Chollet. On the Measure of Intelligence. arXiv:1911.01547 [cs], November 2019. URL http://arxiv.org/abs/1911.01547. arXiv: 1911.01547.
  39. 39.Mathieu Ciancone, Imene Kerboua, Marion Schaeffer, and Wissam Siblini. Extending the massive text embedding benchmark to french, 2024.
  40. 40.cjadams, Daniel Borkan, inversion, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, and nithum. Jigsaw unintended bias in toxicity classification, 2019. URL https://kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification.
  41. 41.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  42. 42.Benjamin Clavié. Jacolbert and hard negatives, towards better japanese-first embeddings for retrieval: Early technical report, 2023.
  43. 43.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems, 2021.
  44. 44.Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld. Specter: Document-level representation learning using citation-informed transformers, 2020a.
  45. 45.Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld. Specter: Document-level representation learning using citation-informed transformers. In ACL, 2020b.
  46. 46.Pierre Colombo, Nathan Noiry, Ekhine Irurozki, and Stephan Clémençon. What are the best systems? new perspectives on nlp benchmarking. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (eds.), Advances in Neural Information Processing Systems, volume 35, pp. 26915–26932. Curran Associates, Inc., 2022. URL https://proceedings.neurips.cc/paper_files/paper/2022/file/ac4920f4085b5662133dd751493946a6-Paper-Conference.pdf.
  47. 47.Tatoeba community. Tatoeba: Collection of sentences and translations, 2021.
  48. 48.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. Xnli: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2018.
  49. 49.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116, 2019.
  50. 50.Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672, 2022.
  51. 51.Benoit Courty, Victor Schmidt, Goyal-Kamal, MarionCoutarel, Boris Feld, Jérémy Lecourt, LiamConnell, SabAmine, inimaz, supatomic, Mathilde Léval, Luis Blanche, Alexis Cruveiller, ouminasara, Franklin Zhao, Aditya Joshi, Alexis Bogroff, Amine Saboni, Hugues de Lavoreille, Niko Laskaris, Edoardo Abati, Douglas Blank, Ziyao Wang, Armin Catovic, alencon, Michał St˛echły, Christian Bauer, Lucas-Otavio, JPW, and MinervaBooks. mlco2/codecarbon: v2.4.1, May 2024. URL https://doi.org/10.5281/zenodo.11171501.
  52. 52.Mathias Creutz. Open subtitles paraphrase corpus for six languages, 2018.
  53. 53.Slawomir Dadas, Michał Perełkiewicz, and Rafał Poswiata. Evaluation of sentence representa- ´ tions in Polish. In Nicoletta Calzolari, Fr’ed’eric B’echet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, H’elène Mazo, Asuncion Moreno, Jan Odijk, and Stelios Piperidis (eds.), Proceedings of the Twelfth Language Resources and Evaluation Conference, pp. 1674–1680, Marseille, France, may 2020. European Language Resources Association. ISBN 979-10-95546-34-4. URL https://aclanthology.org/2020.lrec-1.207.
  54. 54.Sławomir Dadas. Training effective neural sentence encoders from automatically mined paraphrases, 2022.
  55. 55.David Davis. Swahili: News classification dataset (0.2). Zenodo, 2020. doi: 10.5281/zenodo.5514203. URL https://doi.org/10.5281/zenodo.5514203.
  56. 56.Nisansa de Silva. Sinhala text classification: Observations from the perspective of a resource poor language. Year of Publication, 2015.
  57. 57.Leon Derczynski and Alex Speed Kjeldsen. Bornholmsk natural language processing: Resources and tools. In Proceedings of the Nordic Conference of Computational Linguistics (2019), pp. 338–344. Linköping University Electronic Press. URL https://pure.itu.dk/ws/files/84551091/W19_6138.pdf.
  58. 58.Ameet Deshpande, Carlos E Jimenez, Howard Chen, Vishvak Murahari, Victoria Graf, Tanmay Rajpurohit, Ashwin Kalyan, Danqi Chen, and Karthik Narasimhan. Csts: Conditional semantic textual similarity. arXiv preprint arXiv:2305.15093, 2023.
  59. 59.Swapnil Dhanwal, Hritwik Dutta, Hitesh Nankani, Nilay Shrivastava, Yaman Kumar, Junyi Jessy Li, Debanjan Mahata, Rakesh Gosangi, Haimin Zhang, Rajiv Ratn Shah, and Amanda Stent. An annotated dataset of discourse modes in Hindi stories. In Proceedings of the 12th Language Resources and Evaluation Conference, Marseille, France, may 2020. European Language Resources Association. ISBN 979-10-95546-34-4. URL https://www.aclweb.org/anthology/2020.lrec-1.149.
  60. 60.Kaustubh D Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahendiran, Simon Mille, Ashish Shrivastava, Samson Tan, et al. Nl-augmenter: A framework for task-sensitive natural language augmentation. arXiv preprint arXiv:2112.02721, 2021.
  61. 61.Thomas Diggelmann, Jordan Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold. Climate-fever: A dataset for verification of real-world climate claims, 2021.
  62. 62.Kenneth Enevoldsen, Márton Kardos, Niklas Muennighoff, and Kristoffer Laigaard Nielbo. The scandinavian embedding benchmarks: Comprehensive assessment of multilingual and monolingual text embedding. arXiv preprint arXiv:2406.02396, 2024.
  63. 63.Alexander R Fabbri, Wojciech Kry’sci’nski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. Summeval: Re-evaluating summarization evaluation. arXiv preprint arXiv:2007.12626, 2020.
  64. 64.Christian Federmann, Tom Kocmi, and Ying Xin. NTREX-128 – news test references for MT evaluation of 128 languages. In Proceedings of the First Workshop on Scaling Up Multilingual Evaluation, pp. 21–24, Online, nov 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.sumeval-1.4.
  65. 65.Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. Language-agnostic BERT sentence embedding. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 878–891, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.62. URL https://aclanthology.org/2022.acl-long.62.
  66. 66.Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gokhan Tur, and Prem Natarajan. Massive: A 1m-example multilingual natural language understanding dataset with 51 typologically-diverse languages, 2022.
  67. 67.Wikimedia Foundation. Wikimedia downloads. URL https://dumps.wikimedia.org.
  68. 68.Marc Franco-Salvador, Paolo Rosso, and Roberto Navigli. A knowledge-based representation for cross-language document retrieval and categorization. In Shuly Wintner, Sharon Goldwater, and Stefan Riezler (eds.), Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, pp. 414–423, Gothenburg, Sweden, April 2014. Association for Computational Linguistics. doi: 10.3115/v1/E14-1044. URL https://aclanthology.org/E14-1044.
  69. 69.Jay Gala, Pranjal A Chitale, A K Raghavan, Varun Gumma, Sumanth Doddapaneni, Aswanth Kumar M, Janki Atul Nawale, Anupama Sujatha, Ratish Puduppully, Vivek Raghavan, Pratyush Kumar, Mitesh M Khapra, Raj Dabre, and Anoop Kunchukuttan. Indictrans2: Towards high-quality and accessible machine translation models for all 22 scheduled indian languages. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=vfT4YuzAYA.
  70. 70.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework for few-shot language model evaluation, September 2021. URL https://doi.org/10.5281/zenodo.5371628.
  71. 71.Gregor Geigle, Nils Reimers, Andreas R"uckl’e, and Iryna Gurevych. Tweac: Transformer with extendable qa agent classifiers. arXiv preprint, abs/2104.07081, 2021. URL http://arxiv.org/abs/2104.07081.
  72. 72.Tsvetanka Georgieva-Trifonova, Milena Stefanova, and Stefan Kalchev. Dataset for “Customer Feedback Text Analysis for Online Stores Reviews in Bulgarian”, 2018. URL https://doi.org/10.7910/DVN/TXIK9P.
  73. 73.Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. The third PASCAL recognizing textual entailment challenge. In Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, pp. 1–9, Prague, jun 2007. Association for Computational Linguistics. URL https://aclanthology.org/W07-1401.
  74. 74.Hippolyte Gisserot-Boukhlef, Manuel Faysse, Emmanuel Malherbe, Céline Hudelot, and Pierre Colombo. Towards trustworthy reranking: A simple yet effective abstention mechanism. arXiv preprint arXiv:2402.12997, 2024.
  75. 75.Dirk Goldhahn, Thomas Eckart, and Uwe Quasthoff. Building large monolingual dictionaries at the leipzig corpora collection: From 100 to 200 languages. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12), 2012.
  76. 76.Charles J Gomez, Andrew C Herman, and Paolo Parigi. Leading countries in global science increasingly receive more citations than other countries doing similar research. Nature Human Behaviour, 6(7):919–929, 2022.
  77. 77.Matilde González, Clara García, and Lucía Sánchez. Diabla: A corpus of bilingual spontaneous written dialogues for machine translation. In Proceedings of the 12th Language Resources and Evaluation Conference, pp. 4192–4198, 2019.
  78. 78.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, and Francisco Guzm’an. The flores-101 evaluation benchmark for low-resource and multilingual machine translation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 19–35, 2022.
  79. 79.Niklas Gudowsky. Limits and benefits of participatory agenda setting for research and innovation. European Journal of Futures Research, 9(1):8, 2021.
  80. 80.Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N. Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory M. Dickinson, Haggai Porat, Jason Hegland, Jessica Wu, Joe Nudell, Joel Niklaus, John Nay, Jonathan H. Choi, Kevin Tobia, Margaret Hagan, Megan Ma, Michael Livermore, Nikon Rasumov-Rahe, Nils Holzenberger, Noam Kolt, Peter Henderson, Sean Rehaag, Sharad Goel, Shang Gao, Spencer Williams, Sunny Gandhi, Tom Zur, Varun Iyer, and Zehua Li. Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models, 2023.
  81. 81.Ren’e Haas and Leon Derczynski. Discriminating between similar Nordic languages. In Marcos Zampieri, Preslav Nakov, Nikola Ljubesi’c, J"org Tiedemann, Yves Scherrer, and Tommi Jauhiainen (eds.), Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects, pp. 67–75, Kiyv, Ukraine, apr 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.vardial-1.8.
  82. 82.Ivan Habernal, Tom’as Pt’acek, and Josef Steinberger. Sentiment analysis in Czech social media using supervised machine learning. In Alexandra Balahur, Erik van der Goot, and Andres Montoyo (eds.), Proceedings of the 4th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, pp. 65–74, Atlanta, Georgia, jun 2013. Association for Computational Linguistics. URL https://aclanthology.org/W13-1609.
  83. 83.Alexa Hagerty and Igor Rubinov. Global ai ethics: A review of the social impacts and ethical implications of artificial intelligence, 2019.
  84. 84.Margot Hanley, Apoorv Khandelwal, Hadar Averbuch-Elor, Noah Snavely, and Helen Nissenbaum. An ethical highlighter for people-centric dataset creation. arXiv preprint arXiv:2011.13583, 2020.
  85. 85.Mariya Hendriksen, Svitlana Vakulenko, Ernst Kuiper, and Maarten de Rijke. Scene-centric vs. object-centric image-text cross-modal retrieval: A reproducibility study. In European Conference on Information Retrieval, pp. 68–85. Springer, 2023.
  86. 86.Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. Measuring coding challenge competence with apps. NeurIPS, 2021a.
  87. 87.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021b.
  88. 88.Søren Vejlgaard Holm. Are gllms danoliterate? benchmarking generative nlp in danish. 2024.
  89. 89.Doris Hoogeveen, Karin M. Verspoor, and Timothy Baldwin. Cqadupstack: A benchmark data set for community question-answering research. In Proceedings of the 20th Australasian Document Computing Symposium (ADCS), ADCS ’15, pp. 3:1–3:8, New York, NY, USA, 2015. ACM. ISBN 978-1-4503-4040-3. doi: 10.1145/2838931.2838934. URL http://doi.acm.org/10.1145/2838931.2838934.
  90. 90.Christoph Hoppe, David Pelkmann, Nico Migenda, Daniel Hötte, and Wolfram Schenck. Towards intelligent legal advisors for document retrieval and question-answering in german legal documents. In 2021 IEEE Fourth International Conference on Artificial Intelligence and Knowledge Engineering (AIKE), pp. 29–32, 2021. doi: 10.1109/AIKE52691.2021.00011.
  91. 91.Junjie Huang, Duyu Tang, Linjun Shou, Ming Gong, Ke Xu, Daxin Jiang, Ming Zhou, and Nan Duan. Cosqa: 20,000+ web queries for code search and question answering, 2021. URL https://arxiv.org/abs/2105.13239.
  92. 92.Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. CodeSearchNet challenge: Evaluating the state of semantic code search. arXiv preprint arXiv:1909.09436, 2019.
  93. 93.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. Unsupervised dense information retrieval with contrastive learning. arXiv preprint arXiv:2112.09118, 2021.
  94. 94.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7b, 2023.
  95. 95.Dame Jovanoski, Veno Pachovski, and Preslav Nakov. Sentiment analysis in Twitter for Macedonian. In Ruslan Mitkov, Galia Angelova, and Kalina Bontcheva (eds.), Proceedings of the International Conference Recent Advances in Natural Language Processing, pp. 249–257, Hissar, Bulgaria, sep 2015. INCOMA Ltd. Shoumen, BULGARIA. URL https://aclanthology.org/R15-1034.
  96. 96.Ehsan Kamalloo, Aref Jafari, Xinyu Zhang, Nandan Thakur, and Jimmy Lin. HAGRID: A human-llm collaborative dataset for generative information-seeking with attribution. arXiv:2307.16883, 2023.
  97. 97.Jenna Kanerva, Filip Ginter, Li-Hsin Chang, Iiro Rastas, Valtteri Skantsi, Jemina Kilpel"ainen, Hanna-Mari Kupari, Jenna Saarni, Maija Sev’on, and Otto Tarkka. Finnish paraphrase corpus. In Simon Dobnik and Lilja Øvrelid (eds.), Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa), pp. 288–298, Reykjavik, Iceland (Online), 2021. Link"oping University Electronic Press, Sweden. URL https://aclanthology.org/2021.nodalida-main.29.
  98. 98.Jiwon Kim and Won Ik Cho. Kocasm: Korean automatic sarcasm detection. https://github.com/SpellOnYou/korean-sarcasm, 2019.
  99. 99.Anoop Kunchukuttan, Divyanshu Kakwani, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. Ai4bharat-indicnlp corpus: Monolingual corpora and word embeddings for indic languages. arXiv preprint arXiv:2005.00085, 2020.
  100. 100.Tzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein, and Haiyi Zhu. Wikibench: Community-driven data curation for ai evaluation on wikipedia. In Proceedings of the CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400703300. doi: 10.1145/3613904.3642278. URL https://doi.org/10.1145/3613904.3642278.
  101. 101.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. Natural questions: a benchmark for question answering research, 2019. URL https://aclanthology.org/Q19-1026/.
  102. 102.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. Openassistant conversations – democratizing large language model alignment, 2023.
  103. 103.Wuwei Lan, Siyu Qiu, Hua He, and Wei Xu. A continuously growing dataset of sentential paraphrases. In Martha Palmer, Rebecca Hwa, and Sebastian Riedel (eds.), Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 1224–1234, Copenhagen, Denmark, sep 2017. Association for Computational Linguistics. doi: 10.18653/v1/D17-1126. URL https://aclanthology.org/D17-1126.
  104. 104.Ken Lang. Newsweeder: Learning to filter netnews. In Armand Prieditis and Stuart Russell (eds.), Machine Learning Proceedings 1995, pp. 331–339. Morgan Kaufmann, San Francisco (CA), 1995. ISBN 978-1-55860-377-6. doi: https://doi.org/10.1016/B978-1-55860-377-6.50048-7. URL https://www.sciencedirect.com/science/article/pii/B9781558603776500487.
  105. 105.Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. Nv-embed: Improved techniques for training llms as generalist embedding models, 2024.
  106. 106.Jean Lee, Taejun Lim, Heejun Lee, Bogeun Jo, Yangsok Kim, Heegeun Yoon, and Soyeon Caren Han. K-MHaS: A multi-label hate speech detection dataset in Korean online news comment. In Proceedings of the 29th International Conference on Computational Linguistics, pp. 3530–3538, Gyeongju, Republic of Korea, oct 2022. International Committee on Computational Linguistics. URL https://aclanthology.org/2022.coling-1.311.
  107. 107.Antoine Lefebvre-Brossard, Stephane Gazaille, and Michel C. Desmarais. Alloprof: a new french question-answer education dataset and its use in an information retrieval case study, 2023. URL https://arxiv.org/abs/2302.07738.
  108. 108.Joao Augusto Leite, Diego F. Silva, Kalina Bontcheva, and Carolina Scarton. Toxic language detection in social media for brazilian portuguese: New dataset and multilingual analysis. CoRR, abs/2010.04543, 2020. URL https://arxiv.org/abs/2010.04543.
  109. 109.Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. Mlqa: Evaluating cross-lingual extractive question answering. arXiv preprint arXiv:1910.07475, art. arXiv:1910.07475, 2019.
  110. 110.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-augmented generation for knowledge-intensive nlp tasks, 2021. URL https://arxiv.org/abs/2005.11401.
  111. 111.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Sasko, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clement Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander M. Rush, and Thomas Wolf. Datasets: A community library for natural language processing. CoRR, abs/2109.02846, 2021. URL https://arxiv.org/abs/2109.02846.
  112. 112.Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, and Yashar Mehdad. MTOP: A comprehensive multilingual task-oriented semantic parsing benchmark. In Paola Merlo, Jorg Tiedemann, and Reut Tsarfaty (eds.), Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pp. 2950–2962, Online, apr 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.257. URL https://aclanthology.org/2021.eacl-main.257.
  113. 113.Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. Starcoder: may the source be with you!, 2023a.
  114. 114.Xiangyang Li, Kuicai Dong, Yi Quan Lee, Wei Xia, Yichun Yin, Hao Zhang, Yong Liu, Yasheng Wang, and Ruiming Tang. Coir: A comprehensive benchmark for code information retrieval models, 2024. URL https://arxiv.org/abs/2407.02883.
  115. 115.Yudong Li, Yuqing Zhang, Zhe Zhao, Linlin Shen, Weijie Liu, Weiquan Mao, and Hui Zhang. Csl: A large-scale chinese scientific literature dataset, 2022.
  116. 116.Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. Towards general text embeddings with multi-stage contrastive learning, 2023b. URL https://arxiv.org/abs/2308.03281.
  117. 117.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022.
  118. 118.Daniele Licari, Praveen Bushipaka, Gabriele Marino, Giovanni Comand’e, and Tommaso Cucinotta. Legal holding extraction from italian case documents using italian-legal-bert text summarization. In Proceedings of the Nineteenth International Conference on Artificial Intelligence and Law, ICAIL ’23, pp. 148–156, New York, NY, USA, 2023. Association for Computing Machinery. ISBN 9798400701979. doi: 10.1145/3594536.3595177. URL https://doi.org/10.1145/3594536.3595177.
  119. 119.Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang. Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation, 2023.
  120. 120.Shayne Longpre, Robert Mahari, Anthony Chen, Naana Obeng-Marnu, Damien Sileo, William Brannon, Niklas Muennighoff, Nathan Khazam, Jad Kabbara, Kartik Perisetla, Xinyi Wu, Enrico Shippole, Kurt Bollacker, Tongshuang Wu, Luis Villa, Sandy Pentland, Deb Roy, and Sara Hooker. The data provenance initiative: A large scale audit of dataset licensing & attribution in ai, 2023.
  121. 121.Shayne Longpre, Robert Mahari, Ariel Lee, Campbell Lund, Hamidah Oderinwale, William Brannon, Nayan Saxena, Naana Obeng-Marnu, Tobin South, Cole Hunter, Kevin Klyman, Christopher Klamm, Hailey Schoelkopf, Nikhil Singh, Manuel Cherep, Ahmad Anis, An Dinh, Caroline Chitongo, Da Yin, Damien Sileo, Deividas Mataciunas, Diganta Misra, Emad Alghamdi, Enrico Shippole, Jianguo Zhang, Joanna Materzynska, Kun Qian, Kush Tiwary, Lester Miranda, Manan Dey, Minnie Liang, Mohammed Hamdy, Niklas Muennighoff, Seonghyeon Ye, Seungone Kim, Shrestha Mohanty, Vipul Gupta, Vivek Sharma, Vu Minh Chien, Xuhui Zhou, Yizhi Li, Caiming Xiong, Luis Villa, Stella Biderman, Hanlin Li, Daphne Ippolito, Sara Hooker, Jad Kabbara, and Sandy Pentland. Consent in crisis: The rapid decline of the ai data commons, 2024a. URL https://arxiv.org/abs/2407.14933.
  122. 122.Shayne Longpre, Nikhil Singh, Manuel Cherep, Kushagra Tiwary, Joanna Materzynska, William Brannon, Robert Mahari, Manan Dey, Mohammed Hamdy, Nayan Saxena, Ahmad Mustafa Anis, Emad A. Alghamdi, Vu Minh Chien, Naana Obeng-Marnu, Da Yin, Kun Qian, Yizhi Li, Minnie Liang, An Dinh, Shrestha Mohanty, Deividas Mataciunas, Tobin South, Jianguo Zhang, Ariel N. Lee, Campbell S. Lund, Christopher Klamm, Damien Sileo, Diganta Misra, Enrico Shippole, Kevin Klyman, Lester JV Miranda, Niklas Muennighoff, Seonghyeon Ye, Seungone Kim, Vipul Gupta, Vivek Sharma, Xuhui Zhou, Caiming Xiong, Luis Villa, Stella Biderman, Alex Pentland, Sara Hooker, and Jad Kabbara. Bridging the data provenance gap across text, speech and video, 2024b. URL https://arxiv.org/abs/2412.17847.
  123. 123.Holy Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar, Lester James V Miranda, Jennifer Santoso, Elyanah Aco, Akhdan Fadhilah, Jonibek Mansurov, Joseph Marvin Imperial, Onno P Kampman, et al. Seacrowd: A multilingual multimodal data hub and benchmark suite for southeast asian languages. arXiv preprint arXiv:2406.10118, 2024.
  124. 124.Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Zhuang Li, Wen-Ding Li, Megan Risdal, Jia Li, Jian Zhu, Terry Yue Zhuo, Evgenii Zheltonozhskii, Nii Osae Osae Dade, Wenhao Yu, Lucas Krauß, Naman Jain, Yixuan Su, Xuanli He, Manan Dey, Edoardo Abati, Yekun Chai, Niklas Muennighoff, Xiangru Tang, Muhtasham Oblokulov, Christopher Akiki, Marc Marone, Chenghao Mou, Mayank Mishra, Alex Gu, Binyuan Hui, Tri Dao, Armel Zebaze, Olivier Dehaene, Nicolas Patry, Canwen Xu, Julian McAuley, Han Hu, Torsten Scholak, Sebastien Paquet, Jennifer Robinson, Carolyn Jane Anderson, Nicolas Chapados, Mostofa Patwary, Nima Tajbakhsh, Yacine Jernite, Carlos Muñoz Ferrandis, Lingming Zhang, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries. Starcoder 2 and the stack v2: The next generation, 2024. URL https://arxiv.org/abs/2402.19173.
  125. 125.Xing Han Lu, Siva Reddy, and Harm de Vries. The StatCan dialogue dataset: Retrieving data tables through conversations with genuine intents. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pp. 2799–2829, Dubrovnik, Croatia, may 2023. Association for Computational Linguistics. URL https://arxiv.org/abs/2304.01412.
  126. 126.Xing Han Lù, Zdenek Kasner, and Siva Reddy. Weblinx: Real-world website navigation with ˇ multi-turn dialogue, 2024.
  127. 127.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. Learning word vectors for sentiment analysis. In Dekang Lin, Yuji Matsumoto, and Rada Mihalcea (eds.), Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pp. 142–150, Portland, Oregon, USA, jun 2011. Association for Computational Linguistics. URL https://aclanthology.org/P11-1015.
  128. 128.Yash Madhani, Mitesh M. Khapra, and Anoop Kunchukuttan. Bhasa-abhijnaanam: Native-script and romanized language identification for 22 Indic languages. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 816–826, Toronto, Canada, jul 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-short.71. URL https://aclanthology.org/2023.acl-short.71.
  129. 129.Andani Madodonga, Vukosi Marivate, and Matthew Adendorff. Izindaba-tindzaba: Machine learning news categorisation for long and short text for isizulu and siswati. 4, Jan. 2023. doi: 10.55492/dhasa.v4i01.4449. URL https://upjournals.up.ac.za/index.php/dhasa/article/view/4449.
  130. 130.Wei Chen Maggie, Phil Culliton. Tweet sentiment extraction, 2020. URL https://kaggle.com/competitions/tweet-sentiment-extraction.
  131. 131.Rahmad Mahendra, Alham Fikri Aji, Samuel Louvan, Fahrurrozi Rahman, and Clara Vania. IndoNLI: A natural language inference dataset for Indonesian. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 10511–10527, Online and Punta Cana, Dominican Republic, nov 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.emnlp-main.821.
  132. 132.Arthur Malajyan, Karen Avetisyan, and Tsolak Ghukasyan. Arpa: Armenian paraphrase detection corpus and models, 2020.
  133. 133.P. Malo, A. Sinha, P. Korhonen, J. Wallenius, and P. Takala. Good debt or bad debt: Detecting semantic orientations in economic texts. Journal of the Association for Information Science and Technology, 65, 2014.
  134. 134.Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. A SICK cure for the evaluation of compositional distributional semantic models. In Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Hrafn Loftsson, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, and Stelios Piperidis (eds.), Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), pp. 216–223, Reykjavik, Iceland, May 2014. European Language Resources Association (ELRA). URL http://www.lrec-conf.org/proceedings/lrec2014/pdf/363_Paper.pdf.
  135. 135.Vukosi Marivate, Moseli Mots’Oehli, Valencia Wagner, Richard Lastrucci, and Isheanesu Dzingirai. Puoberta: Training and evaluation of a curated language model for setswana. In SACAIR 2023 (To Appear), 2023.
  136. 136.Philip May. Machine translated multilingual sts benchmark dataset. 2021. URL https://github.com/PhilipMay/stsb-multi-mt.
  137. 137.Philip May, Brooke Fujita, and Tom Aarsen. stsb-multi-mt, 2021. URL https://github.com/PhilipMay/stsb-multi-mt. GitHub repository.
  138. 138.Rui Meng, Ye Liu, Shafiq Rayhan Joty, Caiming Xiong, Yingbo Zhou, and Semih Yavuz. Sfr-embedding-mistral:enhance text retrieval with transfer learning. Salesforce AI Research Blog, 2024. URL https://blog.salesforceairesearch.com/sfr-embedded-mistral/.
  139. 139.Yev Meyer, Marjan Emadi, Dhruv Nathawani, Lipika Ramaswamy, Kendrick Boyd, Maarten Van Segbroeck, Matthew Grossman, Piotr Mlocek, and Drew Newberry. Synthetic-Text-To-SQL: A synthetic dataset for training language models to generate sql queries from natural language prompts, April 2024. URL https://huggingface.co/datasets/gretelai/synthetic-text-to-sql.
  140. 140.Roshanak Mirzaee, Hossein Rajaby Faghihi, Qiang Ning, and Parisa Kordjamshidi. Spartqa: A textual question answering benchmark for spatial reasoning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 4582–4598, 2021.
  141. 141.Julius Monsen and Arne J"onsson. A method for building non-english corpora for abstractive text summarization. In Proceedings of CLARIN Annual Conference, 2021.
  142. 142.Niklas Muennighoff. Sgpt: Gpt sentence embeddings for semantic search, 2022.
  143. 143.Niklas Muennighoff, Qian Liu, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro von Werra, and Shayne Longpre. Octopack: Instruction tuning code large language models. arXiv preprint arXiv:2308.07124, 2023a.
  144. 144.Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. MTEB: Massive text embedding benchmark. In Andreas Vlachos and Isabelle Augenstein (eds.), Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pp. 2014–2037, Dubrovnik, Croatia, May 2023b. Association for Computational Linguistics. doi: 10.18653/v1/2023.eacl-main.148. URL https://aclanthology.org/2023.eacl-main.148.
  145. 145.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. Crosslingual generalization through multitask finetuning, 2023c.
  146. 146.Niklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela. Generative representational instruction tuning, 2024.
  147. 147.Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, Nedjma Ousidhoum, David Ifeoluwa Adelani, Seid Muhie Yimam, Ibrahim Sa’id Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Pavel Brazdil, Felermino D’ario M’ario Ant’onio Ali, Davis Davis, Salomey Osei, Bello Shehu Bello, Falalu Ibrahim, Tajuddeen Gwadabe, Samuel Rutunda, Tadesse Belay, Wendimu Baye Messelle, Hailu Beshada Balcha, Sisay Adugna Chala, Hagos Tesfahun Gebremichael, Bernard Opoku, and Steven Arthur. Afrisenti:a twitter sentiment analysis benchmark for african languages. 2023.
  148. 148.Timo Möller, Julian Risch, and Malte Pietsch. Germanquad and germandpr: Improving non-english question answering and passage retrieval, 2021.
  149. 149.Jørgen Johnsen Navjord and Jon-Mikkel Ryen Korsvik. Beyond extractive: advancing abstractive automatic text summarization in norwegian with transformers. Master’s thesis, Norwegian University of Life Sciences, Ås, 2023.
  150. 150.Thong Nguyen, Mariya Hendriksen, Andrew Yates, and Maarten de Rijke. Multimodal learned sparse retrieval with probabilistic expansion control. In European Conference on Information Retrieval, pp. 448–464. Springer, 2024.
  151. 151.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. MS MARCO: A human generated machine reading comprehension dataset. CoRR, abs/1611.09268, 2016. URL http://arxiv.org/abs/1611.09268.
  152. 152.Dan Nielsen. ScandEval: A benchmark for Scandinavian natural language processing. In Tanel Alum"ae and Mark Fishel (eds.), Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa), pp. 185–201, T’orshavn, Faroe Islands, may 2023. University of Tartu Library. URL https://aclanthology.org/2023.nodalida-1.20.
  153. 153.Joel Niklaus, Matthias Stürmer, and Ilias Chalkidis. An empirical study on cross-x transfer for legal judgment prediction, 2022.
  154. 154.Joakim Nivre, Marie-Catherine De Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajic, Christopher D. Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, and Natalia Silveira. Universal dependencies v1: A multilingual treebank collection. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pp. 1659–1666, 2016.
  155. 155.Jeppe Nørregaard and Leon Derczynski. DanFEVER: claim verification dataset for Danish. In Simon Dobnik and Lilja Øvrelid (eds.), Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa), pp. 422–428, Reykjavik, Iceland (Online), 2021. Link"oping University Electronic Press, Sweden. URL https://aclanthology.org/2021.nodalida-main.47.
  156. 156.Maciej Ogrodniczuk and Mateusz Kope’c. The Polish summaries corpus. In Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Hrafn Loftsson, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, and Stelios Piperidis (eds.), Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), pp. 3712–3715, Reykjavik, Iceland, may 2014. European Language Resources Association (ELRA). URL http://www.lrec-conf.org/proceedings/lrec2014/pdf/1211_Paper.pdf.
  157. 157.Maciej Ogrodniczuk and Łukasz Kobylinski (eds.). ´ Proceedings of the PolEval 2019 Workshop, Warsaw, Poland, 2019. Institute of Computer Science, Polish Academy of Sciences. ISBN 978-83-63159-28-3. URL http://2019.poleval.pl/files/poleval2019.pdf.
  158. 158.James O’Neill, Polina Rozenshtein, Ryuichi Kiryo, Motoko Kubota, and Danushka Bollegala. I wish I would have loved this one, but I didn’t – a multilingual dataset for counterfactual detection in product review. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 7092–7108, Online and Punta Cana, Dominican Republic, nov 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.568. URL https://aclanthology.org/2021.emnlp-main.568.
  159. 159.Nedjma Ousidhoum, Shamsuddeen Hassan Muhammad, Mohamed Abdalla, Idris Abdulmumin, Ibrahim Said Ahmad, Sanchit Ahuja, Alham Fikri Aji, Vladimir Araujo, Abinew Ali Ayele, Pavan Baswani, et al. Semrel2024: A collection of semantic textual relatedness datasets for 14 languages. arXiv preprint arXiv:2402.08638, 2024.
  160. 160.Hille Pajupuu, Jaan Pajupuu, Rene Altrov, and Kairi Tamuri. Estonian Valence Corpus / Eesti valentsikorpus. 11 2023. doi: 10.6084/m9.figshare.24517054.v1. URL https://figshare.com/articles/dataset/Estonian_Valence_Corpus_Eesti_valentsikorpus/24517054.
  161. 161.Christos Papaloukas, Ilias Chalkidis, Konstantinos Athinaios, Despina-Athanasia Pantazi, and Manolis Koubarakis. Multi-granular legal topic classification on greek legislation. In Proceedings of the Natural Legal Language Processing Workshop 2021, pp. 63–75, Punta Cana, Dominican Republic, 2021. Association for Computational Linguistics. doi: 10.48550/arXiv.2109.15298. URL https://arxiv.org/abs/2109.15298.
  162. 162.Shantipriya Parida, Sambit Sekhar, Soumendra Kumar Sahoo, Swateek Jena, Abhijeet Parida, Satya Ranjan Dash, and Guneet Singh Kohli. Odiagenai: Generative ai and llm initiative for the odia language. https://huggingface.co/OdiaGenAI, 2023.
  163. 163.Michael Park, Erin Leahey, and Russell J. Funk. Papers and patents are becoming less disruptive over time. Nature, 613:138–144, 2023. URL https://api.semanticscholar.org/CorpusID:255466666.
  164. 164.Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, Adam Khoja, Ryan Kim, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, Daron Anderson, Tung Nguyen, Mobeen Mahmood, Fiona Feng, Steven Y. Feng, Haoran Zhao, Michael Yu, Varun Gangal, Chelsea Zou, Zihan Wang, Jessica P. Wang, Pawan Kumar, Oleksandr Pokutnyi, Robert Gerbicz, Serguei Popov, John-Clark Levin, Mstyslav Kazakov, Johannes Schmitt, Geoff Galgon, Alvaro Sanchez, Yongki Lee, Will Yeadon, Scott Sauers, Marc Roth, Chidozie Agu, Søren Riis, Fabian Giska, Saiteja Utpala, Zachary Giboney, Gashaw M. Goshu, Joan of Arc Xavier, Sarah-Jane Crowson, Mohinder Maheshbhai Naiya, Noah Burns, Lennart Finke, Zerui Cheng, Hyunwoo Park, Francesco Fournier-Facio, John Wydallis, Mark Nandor, Ankit Singh, Tim Gehrunger, Jiaqi Cai, Ben McCarty, Darling Duclosel, Jungbae Nam, Jennifer Zampese, Ryan G. Hoerr, Aras Bacho, Gautier Abou Loume, Abdallah Galal, Hangrui Cao, Alexis C Garretson, Damien Sileo, Qiuyu Ren, Doru Cojoc, Pavel Arkhipov, Usman Qazi, Lianghui Li, Sumeet Motwani, Christian Schroeder de Witt, Edwin Taylor, Johannes Veith, Eric Singer, Taylor D. Hartman, Paolo Rissone, Jaehyeok Jin, Jack Wei Lun Shi, Chris G. Willcocks, Joshua Robinson, Aleksandar Mikov, Ameya Prabhu, Longke Tang, Xavier Alapont, Justine Leon Uro, Kevin Zhou, Emily de Oliveira Santos, Andrey Pupasov Maksimov, Edward Vendrow, Kengo Zenitani, Julien Guillod, Yuqi Li, Joshua Vendrow, Vladyslav Kuchkin, Ng Ze-An, Pierre Marion, Denis Efremov, Jayson Lynch, Kaiqu Liang, Andrew Gritsevskiy, Dakotah Martinez, Ben Pageler, Nick Crispino, Dimitri Zvonkine, Natanael Wildner Fraga, Saeed Soori, Ori Press, Henry Tang, Julian Salazar, Sean R. Green, Lina Brüssel, Moon Twayana, Aymeric Dieuleveut, T. Ryan Rogers, Wenjin Zhang, Bikun Li, Jinzhou Yang, Arun Rao, Gabriel Loiseau, Mikhail Kalinin, Marco Lukas, Ciprian Manolescu, Subrata Mishra, Ariel Ghislain Kemogne Kamdoum, Tobias Kreiman, Tad Hogg, Alvin Jin, Carlo Bosio, Gongbo Sun, Brian P Coppola, Tim Tarver, Haline Heidinger, Rafael Sayous, Stefan Ivanov, Joseph M Cavanagh, Jiawei Shen, Joseph Marvin Imperial, Philippe Schwaller, Shaipranesh Senthilkuma, Andres M Bran, Ali Dehghan, Andres Algaba, Brecht Verbeken, David Noever, Ragavendran P V, Lisa Schut, Ilia Sucholutsky, Evgenii Zheltonozhskii, Derek Lim, Richard Stanley, Shankar Sivarajan, Tong Yang, John Maar, Julian Wykowski, Martí Oller, Jennifer Sandlin, Anmol Sahu, Yuzheng Hu, Sara Fish, Nasser Heydari, Archimedes Apronti, Kaivalya Rawal, Tobias Garcia Vilchis, Yuexuan Zu, Martin Lackner, James Koppel, Jeremy Nguyen, Daniil S. Antonenko, Steffi Chern, Bingchen Zhao, Pierrot Arsene, Alan Goldfarb, Sergey Ivanov, Rafał Poswiata, Chenguang Wang, Daofeng Li, ´ Donato Crisostomi, Andrea Achilleos, Benjamin Myklebust, Archan Sen, David Perrella, Nurdin Kaparov, Mark H Inlow, Allen Zang, Elliott Thornley, Daniil Orel, Vladislav Poritski, Shalev Ben-David, Zachary Berger, Parker Whitfill, Michael Foster, Daniel Munro, Linh Ho, Dan Bar Hava, Aleksey Kuchkin, Robert Lauff, David Holmes, Frank Sommerhage, Keith Schneider, Zakayo Kazibwe, Nate Stambaugh, Mukhwinder Singh, Ilias Magoulas, Don Clarke, Dae Hyun Kim, Felipe Meneguitti Dias, Veit Elser, Kanu Priya Agarwal, Victor Efren Guadarrama Vilchis, Immo Klose, Christoph Demian, Ujjwala Anantheswaran, Adam Zweiger, Guglielmo Albani, Jeffery Li, Nicolas Daans, Maksim Radionov, Václav Rozhon, Ziqiao Ma, Christian Stump, ˇ Mohammed Berkani, Jacob Platnick, Volodymyr Nevirkovets, Luke Basler, Marco Piccardo, Ferenc Jeanplong, Niv Cohen, Josef Tkadlec, Paul Rosu, Piotr Padlewski, Stanislaw Barzowski, Kyle Montgomery, Aline Menezes, Arkil Patel, Zixuan Wang, Jamie Tucker-Foltz, Jack Stade, Tom Goertzen, Fereshteh Kazemi, Jeremiah Milbauer, John Arnold Ambay, Abhishek Shukla, Yan Carlos Leyva Labrador, Alan Givré, Hew Wolff, Vivien Rossbach, Muhammad Fayez Aziz, Younesse Kaddar, Yanxu Chen, Robin Zhang, Jiayi Pan, Antonio Terpin, Niklas Muennighoff, Hailey Schoelkopf, Eric Zheng, Avishy Carmi, Adam Jones, Jainam Shah, Ethan D. L. Brown, Kelin Zhu, Max Bartolo, Richard Wheeler, Andrew Ho, Shaul Barkan, Jiaqi Wang, Martin Stehberger, Egor Kretov, Kaustubh Sridhar, Zienab EL-Wasif, Anji Zhang, Daniel Pyda, Joanna Tam, David M. Cunningham, Vladimir Goryachev, Demosthenes Patramanis, Michael Krause, Andrew Redenti, Daniel Bugas, David Aldous, Jesyin Lai, Shannon Coleman, Mohsen Bahaloo, Jiangnan Xu, Sangwon Lee, Sandy Zhao, Ning Tang, Michael K. Cohen, Micah Carroll, Orr Paradise, Jan Hendrik Kirchner, Stefan Steinerberger, Maksym Ovchynnikov, Jason O. Matos, Adithya Shenoy, Benedito Alves de Oliveira Junior, Michael Wang, Yuzhou Nie, Paolo Giordano, Philipp Petersen, Anna Sztyber-Betley, Priti Shukla, Jonathan Crozier, Antonella Pinto, Shreyas Verma, Prashant Joshi, Zheng-Xin Yong, Allison Tee, Jérémy Andréoletti, Orion Weller, Raghav Singhal, Gang Zhang, Alexander Ivanov, Seri Khoury, Hamid Mostaghimi, Kunvar Thaman, Qijia Chen, Tran Quoc Khánh, Jacob Loader, Stefano Cavalleri, Hannah Szlyk, Zachary Brown, Jonathan Roberts, William Alley, Kunyang Sun, Ryan Stendall, Max Lamparth, Anka Reuel, Ting Wang, Hanmeng Xu, Sreenivas Goud Raparthi, Pablo Hernández-Cámara, Freddie Martin, Dmitry Malishev, Thomas Preu, Tomek Korbak, Marcus Abramovitch, Dominic Williamson, Ziye Chen, Biró Bálint, M Saiful Bari, Peyman Kassani, Zihao Wang, Behzad Ansarinejad, Laxman Prasad Goswami, Yewen Sun, Hossam Elgnainy, Daniel Tordera, George Balabanian, Earth Anderson, Lynna Kvistad, Alejandro José Moyano, Rajat Maheshwari, Ahmad Sakor, Murat Eron, Isaac C. McAlister, Javier Gimenez, Innocent Enyekwe, Andrew Favre D. O., Shailesh Shah, Xiaoxiang Zhou, Firuz Kamalov, Ronald Clark, Sherwin Abdoli, Tim Santens, Khalida Meer, Harrison K Wang, Kalyan Ramakrishnan, Evan Chen, Alessandro Tomasiello, G. Bruno De Luca, Shi-Zhuo Looi, Vinh-Kha Le, Noam Kolt, Niels Mündler, Avi Semler, Emma Rodman, Jacob Drori, Carl J Fossum, Milind Jagota, Ronak Pradeep, Honglu Fan, Tej Shah, Jonathan Eicher, Michael Chen, Kushal Thaman, William Merrill, Carter Harris, Jason Gross, Ilya Gusev, Asankhaya Sharma, Shashank Agnihotri, Pavel Zhelnov, Siranut Usawasutsakorn, Mohammadreza Mofayezi, Sergei Bogdanov, Alexander Piperski, Marc Carauleanu, David K. Zhang, Dylan Ler, Roman Leventov, Ignat Soroko, Thorben Jansen, Pascal Lauer, Joshua Duersch, Vage Taamazyan, Wiktor Morak, Wenjie Ma, William Held, Tran Ðuc Huy, Ruicheng Xian, Armel Randy Zebaze, Mohanad Mohamed, Julian Noah Leser, Michelle X Yuan, Laila Yacar, Johannes Lengler, Hossein Shahrtash, Edson Oliveira, Joseph W. Jackson, Daniel Espinosa Gonzalez, Andy Zou, Muthu Chidambaram, Timothy Manik, Hector Haffenden, Dashiell Stander, Ali Dasouqi, Alexander Shen, Emilien Duc, Bita Golshani, David Stap, Mikalai Uzhou, Alina Borisovna Zhidkovskaya, Lukas Lewark, Mátyás Vincze, Dustin Wehr, Colin Tang, Zaki Hossain, Shaun Phillips, Jiang Muzhen, Fredrik Ekström, Angela Hammon, Oam Patel, Nicolas Remy, Faraz Farhidi, George Medley, Forough Mohammadzadeh, Madellene Peñaflor, Haile Kassahun, Alena Friedrich, Claire Sparrow, Taom Sakal, Omkar Dhamane, Ali Khajegili Mirabadi, Eric Hallman, Mike Battaglia, Mohammad Maghsoudimehrabani, Hieu Hoang, Alon Amit, Dave Hulbert, Roberto Pereira, Simon Weber, Stephen Mensah, Nathan Andre, Anton Peristyy, Chris Harjadi, Himanshu Gupta, Stephen Malina, Samuel Albanie, Will Cai, Mustafa Mehkary, Frank Reidegeld, Anna-Katharina Dick, Cary Friday, Jasdeep Sidhu, Wanyoung Kim, Mariana Costa, Hubeyb Gurdogan, Brian Weber, Harsh Kumar, Tong Jiang, Arunim Agarwal, Chiara Ceconello, Warren S. Vaz, Chao Zhuang, Haon Park, Andrew R. Tawfeek, Daattavya Aggarwal, Michael Kirchhof, Linjie Dai, Evan Kim, Johan Ferret, Yuzhou Wang, Minghao Yan, Krzysztof Burdzy, Lixin Zhang, Antonio Franca, Diana T. Pham, Kang Yong Loh, Joshua Robinson, Shreen Gul, Gunjan Chhablani, Zhehang Du, Adrian Cosma, Colin White, Robin Riblet, Prajvi Saxena, Jacob Votava, Vladimir Vinnikov, Ethan Delaney, Shiv Halasyamani, Syed M. Shahid, Jean-Christophe Mourrat, Lavr Vetoshkin, Renas Bacho, Vincent Ginis, Aleksandr Maksapetyan, Florencia de la Rosa, Xiuyu Li, Guillaume Malod, Leon Lang, Julien Laurendeau, Fatimah Adesanya, Julien Portier, Lawrence Hollom, Victor Souza, Yuchen Anna Zhou, Yigit Yalın, Gbenga Daniel Obikoya, Luca Arnaboldi, Rai, Filippo Bigi, ˘ Kaniuar Bacho, Pierre Clavier, Gabriel Recchia, Mara Popescu, Nikita Shulga, Ngefor Mildred Tanwie, Thomas C. H. Lux, Ben Rank, Colin Ni, Alesia Yakimchyk, Huanxu, Liu, Olle Häggström, Emil Verkama, Himanshu Narayan, Hans Gundlach, Leonor Brito-Santana, Brian Amaro, Vivek Vajipey, Rynaa Grover, Yiyang Fan, Gabriel Poesia Reis e Silva, Linwei Xin, Yosi Kratish, Jakub Łucki, Wen-Ding Li, Justin Xu, Kevin Joseph Scaria, Freddie Vargus, Farzad Habibi, Long, Lian, Emanuele Rodolà, Jules Robins, Vincent Cheng, Declan Grabb, Ida Bosio, Tony Fruhauff, Ido Akov, Eve J. Y. Lo, Hao Qi, Xi Jiang, Ben Segev, Jingxuan Fan, Sarah Martinson, Erik Y. Wang, Kaylie Hausknecht, Michael P. Brenner, Mao Mao, Yibo Jiang, Xinyu Zhang, David Avagian, Eshawn Jessica Scipio, Muhammad Rehan Siddiqi, Alon Ragoler, Justin Tan, Deepakkumar Patil, Rebeka Plecnik, Aaron Kirtland, Roselynn Grace Montecillo, Stephane Durand, Omer Faruk Bodur, Zahra Adoul, Mohamed Zekry, Guillaume Douville, Ali Karakoc, Tania C. B. Santos, Samir Shamseldeen, Loukmane Karim, Anna Liakhovitskaia, Nate Resman, Nicholas Farina, Juan Carlos Gonzalez, Gabe Maayan, Sarah Hoback, Rodrigo De Oliveira Pena, Glen Sherman, Hodjat Mariji, Rasoul Pouriamanesh, Wentao Wu, Gözdenur Demir, Sandra Mendoza, Ismail Alarab, Joshua Cole, Danyelle Ferreira, Bryan Johnson, Hsiaoyun Milliron, Mohammad Safdari, Liangti Dai, Siriphan Arthornthurasuk, Alexey Pronin, Jing Fan, Angel Ramirez-Trinidad, Ashley Cartwright, Daphiny Pottmaier, Omid Taheri, David Outevsky, Stanley Stepanic, Samuel Perry, Luke Askew, Raúl Adrián Huerta Rodríguez, Abdelkader Dendane, Sam Ali, Ricardo Lorena, Krishnamurthy Iyer, Sk Md Salauddin, Murat Islam, Juan Gonzalez, Josh Ducey, Russell Campbell, Maja Somrak, Vasilios Mavroudis, Eric Vergo, Juehang Qin, Benjámin Borbás, Eric Chu, Jack Lindsey, Anil Radhakrishnan, Antoine Jallon, I. M. J. McInnis, Alex Hoover, Sören Möller, Song Bian, John Lai, Tejal Patwardhan, Summer Yue, Alexandr Wang, and Dan Hendrycks. Humanity’s last exam, 2025. URL https://arxiv.org/abs/2501.14249.
  165. 165.Lidia Pivovarova, Ekaterina Pronoza, Elena Yagunova, and Anton Pronoza. Paraphraser: Russian paraphrase corpus and shared task. In Conference on artificial intelligence and natural language, pp. 211–225. Springer, 2017.
  166. 166.Rafał Poswiata, Sławomir Dadas, and Michał Perełkiewicz. PL-MTEB: Polish Massive Text Embed- ´ ding Benchmark. arXiv preprint arXiv:2405.10138, 2024.
  167. 167.Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kakwani, Navneet Kumar, Aswin Pradeep, Srihari Nagaraj, Kumar Deepak, Vivek Raghavan, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh Shantadevi Khapra. Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages. Transactions of the Association for Computational Linguistics, 10:145–162, 02 2022. ISSN 2307-387X. doi: 10.1162/tacl_a_00452. URL https://doi.org/10.1162/tacl_a_00452.
  168. 168.Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks, 2019.
  169. 169.Nils Reimers, Philip Beyer, and Iryna Gurevych. Task-oriented intrinsic evaluation of semantic textual similarity. In Yuji Matsumoto and Rashmi Prasad (eds.), Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pp. 87–96, Osaka, Japan, December 2016. The COLING 2016 Organizing Committee. URL https://aclanthology.org/C16-1009.
  170. 170.Neil Christian R. Riego, Danny Bell Villarba, Ariel Antwaun Rolando C. Sison, Fernandez C. Pineda, and Herminiño C. Lagunzad. Enhancement to low-resource text classification via sequential transfer learning. United International Journal for Research and Technology, 04:72–82, 2023.
  171. 171.Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen Voorhees, Lucy Lu Wang, and William R Hersh. Searching for scientific evidence in a pandemic: An overview of trec-covid, 2021.
  172. 172.Anna Rogers, Olga Kovaleva, Matthew Downey, and Anna Rumshisky. Getting closer to ai complete question answering: A set of prerequisite real tasks. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pp. 8722–8731, 2020.
  173. 173.Andrew Rosenberg and Julia Hirschberg. V-measure: A conditional entropy-based external cluster evaluation measure. In Jason Eisner (ed.), Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL), pp. 410–420, Prague, Czech Republic, June 2007. Association for Computational Linguistics. URL https://aclanthology.org/D07-1043.
  174. 174.Paul R"ottger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert. HateCheck: Functional tests for hate speech detection models. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 41–58, Online, aug 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.4. URL https://aclanthology.org/2021.acl-long.4.
  175. 175.Ivan Rybin, Vladislav Korablinov, Pavel Efimov, and Pavel Braslavski. Rubq 2.0: An innovated russian question answering dataset. In ESWC, pp. 532–547, 2021.
  176. 176.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  177. 177.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. Social iqa: Commonsense reasoning about social interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 4463–4473, 2019.
  178. 178.Salim Sazzed. Cross-lingual sentiment classification in low-resource bengali language. In Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020), pp. 50–60, 2020.
  179. 179.Alexander Sboev, Aleksandr Naumov, and Roman Rybka. Data-driven model for emotion detection in russian texts. Procedia Computer Science, 190:637–642, 2021.
  180. 180.Konstantinos Sechidis, Grigorios Tsoumakas, and Ioannis Vlahavas. On the stratification of multilabel data. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2011, Athens, Greece, September 5-9, 2011, Proceedings, Part III 22, pp. 145–158. Springer, 2011.
  181. 181.Darsh Shah, Tao Lei, Alessandro Moschitti, Salvatore Romeo, and Preslav Nakov. Adversarial domain adaptation for duplicate question detection. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 1056–1063, Brussels, Belgium, 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1131. URL https://aclanthology.org/D18-1131.
  182. 182.Zareen Sharf. Roman Urdu Data Set. UCI Machine Learning Repository, 2018. DOI: https://doi.org/10.24432/C58325.
  183. 183.Eva Sharma, Chen Li, and Lu Wang. BIGPATENT: A large-scale dataset for abstractive and coherent summarization. CoRR, abs/1906.03741, 2019. URL http://arxiv.org/abs/1906.03741.
  184. 184.Tatiana Shavrina, Alena Fenogenova, Anton Emelyanov, Denis Shevelev, Ekaterina Artemova, Valentin Malykh, Vladislav Mikhailov, Maria Tikhonova, Andrey Chertok, and Andrey Evlampiev. Russiansuperglue: A russian language understanding evaluation benchmark. arXiv preprint arXiv:2010.15925, 2020.
  185. 185.Emily Sheng and David Uthus. Investigating societal biases in a poetry composition system, 2020.
  186. 186.Iyanuoluwa Shode, David Ifeoluwa Adelani, Jing Peng, and Anna Feldman. Nollysenti: Leveraging transfer learning and machine translation for nigerian movie sentiment classification. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 986–998, 2023.
  187. 187.Harman Singh, Nitish Gupta, Shikhar Bharadwaj, Dinesh Tewari, and Partha Talukdar. Indicgenbench: A multilingual benchmark to evaluate generation capabilities of llms on indic languages, 2024a.
  188. 188.Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzeminski, ´ Hakimeh Fadaei, Irem Ergün, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Vu Minh Chien, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet Üstün, Marzieh Fadaee, and Sara Hooker. Aya dataset: An open-access collection for multilingual instruction tuning, 2024b.
  189. 189.Artem Snegirev, Maria Tikhonova, Anna Maksimova, Alena Fenogenova, and Alexander Abramov. The russian-focused embedders’ exploration: rumteb benchmark and russian embedding model design, 2024. URL https://arxiv.org/abs/2408.12503.
  190. 190.Vésteinn Snæbjarnarson, Annika Simonsen, Goran Glavaš, and Ivan Vulic. Transfer to a low-resource ´ language via close relatives: The case study on faroese. In Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa), Tórshavn, Faroe Islands, may 22–24 2023. Link"oping University Electronic Press, Sweden.
  191. 191.Ian Soboroff and Stephen Robertson. Building a filtering test collection for trec 2002. In Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval, pp. 243–250, 2003.
  192. 192.Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. Mpnet: Masked and permuted pre-training for language understanding. Advances in neural information processing systems, 33: 16857–16867, 2020.
  193. 193.Gizem Sogancıo ˘ glu, Hakime Öztürk, and Arzucan Özgür. BIOSSES: a semantic sentence similarity ˘ estimation system for the biomedical domain. Bioinformatics, 33(14):i49–i58, 07 2017. ISSN 1367-4803. doi: 10.1093/bioinformatics/btx238. URL https://doi.org/10.1093/bioinformatics/btx238.
  194. 194.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2023.
  195. 195.Michal Stef’anik, Marek Kadlc’ık, Piotr Gramacki, and Petr Sojka. Resources and few-shot learners for in-context learning in slavic languages. arXiv preprint arXiv:2304.01922, 2023.
  196. 196.Hongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi, Niklas Muennighoff, Han yu Wang, Haisu Liu, Quan Shi, Zachary S. Siegel, Michael Tang, Ruoxi Sun, Jinsung Yoon, Sercan O. Arik, Danqi Chen, and Tao Yu. Bright: A realistic and challenging benchmark for reasoning-intensive retrieval, 2024. URL https://arxiv.org/abs/2407.12883.
  197. 197.Piotr Szymanski and Tomasz Kajdanowicz. A network perspective on stratification of multi- ´ label data. In Paula Branco Luís Torgo and Nuno Moniz (eds.), Proceedings of the First International Workshop on Learning with Imbalanced Domains: Theory and Applications, volume 74 of Proceedings of Machine Learning Research, pp. 22–35. PMLR, 22 Sep 2017. URL https://proceedings.mlr.press/v74/szyma%C5%84ski17a.html.
  198. 198.Qingyu Tan, Hwee Tou Ng, and Lidong Bing. Towards benchmarking and improving the temporal reasoning capability of large language models. arXiv preprint arXiv:2306.08952, 2023.
  199. 199.Nandan Thakur, Nils Reimers, Andreas R"uckl’e, Abhishek Srivastava, and Iryna Gurevych. BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://openreview.net/forum?id=wCu6T5xFjeJ.
  200. 200."Nandan Thakur, Luiz Bonifacio, Maik Fr"obe, Alexander Bondarenko, Ehsan Kamalloo, Martin Potthast, Matthias Hagen, and Jimmy Lin". "systematic evaluation of neural retrieval models on the Touch’e 2020 argument retrieval subset of BEIR". In "Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval", 2024.
  201. 201.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. FEVER: a large-scale dataset for fact extraction and VERification. In Marilyn Walker, Heng Ji, and Amanda Stent (eds.), Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 809–819, New Orleans, Louisiana, jun 2018a. Association for Computational Linguistics. doi: 10.18653/v1/N18-1074. URL https://aclanthology.org/N18-1074.
  202. 202.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. Fever: a large-scale dataset for fact extraction and verification, 2018b.
  203. 203.J"org Tiedemann and Santhosh Thottingal. Opus-mt — building open translation services for the world. In Proceedings of the 22nd Annual Conference of the European Association for Machine Translation (EAMT), 2020.
  204. 204.Herbert Ullrich, Jan Drchal, Martin Rypar, Hana Vincourov’a, and V’aclav Moravec. Csfever and ` ctkfacts: acquiring czech data for fact verification. Language Resources and Evaluation, 57(4): 1571–1605, 2023.
  205. 205.Elena Volodina, Yousuf Ali Mohammed, and Julia Klezl. Dalaj - a dataset for linguistic acceptability judgments for swedish: Format, baseline, sharing, 2021.
  206. 206.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Tal Linzen, Grzegorz Chrupała, and Afra Alishahi (eds.), Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pp. 353–355, Brussels, Belgium, November 2018. Association for Computational Linguistics. doi: 10.18653/v1/W18-5446. URL https://aclanthology.org/W18-5446.
  207. 207.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. Superglue: A stickier benchmark for general-purpose language understanding systems. CoRR, abs/1905.00537, 2019. URL http://arxiv.org/abs/1905.00537.
  208. 208.Kexin Wang, Nils Reimers, and Iryna Gurevych. Tsdae: Using transformer-based sequential denoising auto-encoderfor unsupervised sentence embedding learning. arXiv preprint arXiv:2104.06979, 4 2021a. URL https://arxiv.org/abs/2104.06979.
  209. 209.Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. Text embeddings by weakly-supervised contrastive pre-training, 2022.
  210. 210.Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. Improving text embeddings with large language models. arXiv preprint arXiv:2401.00368, 2023.
  211. 211.Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. Multilingual e5 text embeddings: A technical report. arXiv preprint arXiv:2402.05672, 2024.
  212. 212.Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei. MiniLMv2: Multi-head self-attention relation distillation for compressing pretrained transformers. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (eds.), Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pp. 2140–2151, Online, August 2021b. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-acl.188. URL https://aclanthology.org/2021.findings-acl.188.
  213. 213.Silvan Wehrli, Bert Arnrich, and Christopher Irrgang. German text embedding clustering benchmark, 2024. URL https://arxiv.org/abs/2401.02709.
  214. 214.Orion Weller, Benjamin Chang, Sean MacAvaney, Kyle Lo, Arman Cohan, Benjamin Van Durme, Dawn Lawrie, and Luca Soldaini. Followir: Evaluating and teaching information retrieval models to follow instructions, 2024.
  215. 215.Orion Weller, Benjamin Chang, Eugene Yang, Mahsa Yarmohammadi, Sam Barham, Sean MacAvaney, Arman Cohan, Luca Soldaini, Benjamin Van Durme, and Dawn Lawrie. mfollowir: a multilingual benchmark for instruction following in retrieval, 2025. URL https://arxiv.org/abs/2501.19264.
  216. 216.Andika William and Yunita Sari. Click-id: A novel dataset for indonesian clickbait headlines. Data in Brief, 32:106231, 2020. ISSN 2352-3409. doi: https://doi.org/10.1016/j.dib.2020.106231. URL http://www.sciencedirect.com/science/article/pii/S2352340920311252.
  217. 217.Genta Winata, Lingjue Xie, Karthik Radhakrishnan, Yifan Gao, and Daniel Preo¸tiuc-Pietro. Efficient zero-shot cross-lingual inference via retrieval. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 93–104, 2023a.
  218. 218.Genta Indra Winata, Andrea Madotto, Zhaojiang Lin, Rosanne Liu, Jason Yosinski, and Pascale Fung. Language models are few-shot multilingual learners. In Proceedings of the 1st Workshop on Multilingual Representation Learning, pp. 1–15, 2021.
  219. 219.Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, and Sebastian Ruder. Nusax: Multilingual parallel sentiment dataset for 10 indonesian local languages, 2022.
  220. 220.Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, et al. Nusax: Multilingual parallel sentiment dataset for 10 indonesian local languages. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pp. 815–834, 2023b.
  221. 221.Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Yutong Wang, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, et al. Worldcuisines: A massive-scale benchmark for multilingual and multicultural visual question answering on global cuisines. arXiv preprint arXiv:2410.12705, 2024a.
  222. 222.Genta Indra Winata, Ruochen Zhang, and David Ifeoluwa Adelani. Miners: Multilingual language models as semantic retrievers. arXiv preprint arXiv:2406.07424, 2024b.
  223. 223.Marco Wrzalik and Dirk Krechel. GerDaLIR: A German dataset for legal information retrieval. In Proceedings of the Natural Legal Language Processing Workshop 2021, pp. 123–128, Punta Cana, Dominican Republic, nov 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.nllp-1.13.
  224. 224.Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, and Ming Zhou. MIND: A large-scale dataset for news recommendation. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 3597–3606, Online, jul 2020a. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.331. URL https://aclanthology.org/2020.acl-main.331.
  225. 225.Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, et al. Mind: A large-scale dataset for news recommendation. In Proceedings of the 58th annual meeting of the association for computational linguistics, pp. 3597–3606, 2020b.
  226. 226.Mengzhou Xia, Antonios Anastasopoulos, Ruochen Xu, Yiming Yang, and Graham Neubig. Predicting performance for natural language processing tasks. CoRR, abs/2005.00870, 2020. URL https://arxiv.org/abs/2005.00870.
  227. 227.Chenghao Xiao, G Thomas Hudson, and Noura Al Moubayed. Rar-b: Reasoning as retrieval benchmark. arXiv preprint arXiv:2404.06347, 2024a.
  228. 228.Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. C-pack: Packaged resources to advance general chinese embedding, 2023.
  229. 229.Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. C-pack: Packed resources for general chinese embeddings. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 641–649, 2024b.
  230. 230.Xiaohui Xie, Qian Dong, Bingning Wang, Feiyang Lv, Ting Yao, Weinan Gan, Zhijing Wu, Xiangsheng Li, Haitao Li, Yiqun Liu, and Jin Ma. T2ranking: A large-scale chinese benchmark for passage ranking, 2023.
  231. 231.Wei Xu, Chris Callison-Burch, and Bill Dolan. SemEval-2015 task 1: Paraphrase and semantic similarity in Twitter (PIT). In Preslav Nakov, Torsten Zesch, Daniel Cer, and David Jurgens (eds.), Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval 2015), pp. 1–11, Denver, Colorado, jun 2015. Association for Computational Linguistics. doi: 10.18653/v1/S15-2001. URL https://aclanthology.org/S15-2001.
  232. 232.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934, 2020.
  233. 233.Weixiang Yan, Yuchen Tian, Yunzhe Li, Qian Chen, and Wen Wang. Codetransocean: A comprehensive multilingual benchmark for code translation, 2023. URL https://arxiv.org/abs/2310.04951.
  234. 234.Hitomi Yanaka and Koji Mineshima. Compositional evaluation on japanese textual entailment and similarity. Transactions of the Association for Computational Linguistics, 10:1266–1284, 2022.
  235. 235.Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. Paws-x: A cross-lingual adversarial dataset for paraphrase identification, 2019.
  236. 236.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 2369–2380, Brussels, Belgium, 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1259. URL https://aclanthology.org/D18-1259.
  237. 237.Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. Metamath: Bootstrap your own mathematical questions for large language models. arXiv preprint arXiv:2309.12284, 2023.
  238. 238.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.
  239. 239.Xiang Zhang, Junbo Zhao, and Yann LeCun. Character-level convolutional networks for text classification. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett (eds.), Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc., 2015. URL https://proceedings.neurips.cc/paper_files/paper/2015/file/250cf8b51c773f3f8dc8b4be867a9a02-Paper.pdf.
  240. 240.Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin. Making a miracl: Multilingual information retrieval across a continuum of languages, 2022.
  241. 241.Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin. MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse Languages. Transactions of the Association for Computational Linguistics, 11:1114–1131, 09 2023. ISSN 2307-387X. doi: 10.1162/tacl_a_00595. URL https://doi.org/10.1162/tacl_a_00595.
  242. 242.Tianyu Zheng, Ge Zhang, Tianhao Shen, Xueling Liu, Bill Yuchen Lin, Jie Fu, Wenhu Chen, and Xiang Yue. Opencodeinterpreter: Integrating code generation with execution and refinement. arXiv preprint arXiv:2402.14658, 2024.
  243. 243.Dawei Zhu, Liang Wang, Nan Yang, Yifan Song, Wenhao Wu, Furu Wei, and Sujian Li. Longembed: Extending embedding models for long context retrieval. arXiv preprint arXiv:2404.12096, 2024.
  244. 244.Elena Zotova, Rodrigo Agerri, Manuel Nuñez, and German Rigau. Multilingual stance detection in tweets: The Catalonia independence corpus. In Nicoletta Calzolari, Fr’ed’eric B’echet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, H’elène Mazo, Asuncion Moreno, Jan Odijk, and Stelios Piperidis (eds.), Proceedings of the Twelfth Language Resources and Evaluation Conference, pp. 1368–1375. European Language Resources Association, may 2020. ISBN 979-10-95546-34-4.
  245. 245.Pierre Zweigenbaum, Serge Sharoff, and Reinhard Rapp. Overview of the second BUCC shared task: Spotting parallel sentences in comparable corpora. In Serge Sharoff, Pierre Zweigenbaum, and Reinhard Rapp (eds.), Proceedings of the 10th Workshop on Building and Using Comparable Corpora, pp. 60–67, Vancouver, Canada, aug 2017. Association for Computational Linguistics. doi: 10.18653/v1/W17-2512. URL https://aclanthology.org/W17-2512.
  246. 246.Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. Aya model: An instruction finetuned open-access multilingual language model, 2024.
  247. 247.Łukasz Augustyniak, Kamil Tagowski, Albert Sawczyn, Denis Janiak, Roman Bartusiak, Adrian Szymczak, Marcin W ˛atroba, Arkadiusz Janz, Piotr Szymanski, Mikołaj Morzy, Tomasz Kajdanow- ´ icz, and Maciej Piasecki. This is the way: designing and compiling lepiszcze, a comprehensive nlp benchmark for polish, 2022. URL https://arxiv.org/abs/2211.13112.
  248. 248.Michal Štefánik, Marek Kadlcík, Piotr Gramacki, and Petr Sojka. Resources and few-shot learners ˇ for in-context learning in slavic languages, 2023.

Citation

MLA
Enevoldsen, K., et al. “MMTEB: Massive Multilingual Text Embedding Benchmark”. arXiv, 2025, http://arxiv.org/abs/2502.13595v4.
APA
Enevoldsen, K., Chung, I., Kerboua, I., Kardos, M., Mathur, A., Stap, D., Gala, J., Siblini, W., Krzemiński, D., Winata, G. I., Sturua, S., Utpala, S., Ciancone, M., Schaeffer, M., Sequeira, G., Misra, D., Dhakal, S., Rystrøm, J., Solomatin, R., … Muennighoff, N. (2025). MMTEB: Massive Multilingual Text Embedding Benchmark. arXiv. http://arxiv.org/abs/2502.13595v4
Chicago
Enevoldsen, K., I. Chung, I. Kerboua, et al. 2025. “MMTEB: Massive Multilingual Text Embedding Benchmark”. arXiv. http://arxiv.org/abs/2502.13595v4.
Harvard
Enevoldsen, K. et al. (2025) “MMTEB: Massive Multilingual Text Embedding Benchmark”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.13595v4.
Vancouver
1. Enevoldsen K, Chung I, Kerboua I, et al (2025) MMTEB: Massive Multilingual Text Embedding Benchmark. arXiv

BibTeX

@article{enevoldsen2025mmteb,
  title = {MMTEB: Massive Multilingual Text Embedding Benchmark},
  author = {Enevoldsen, Kenneth and Chung, Isaac and Kerboua, Imene and Kardos, Márton and Mathur, Ashwin and Stap, David and Gala, Jay and Siblini, Wissam and Krzemiński, Dominik and Winata, Genta Indra and Sturua, Saba and Utpala, Saiteja and Ciancone, Mathieu and Schaeffer, Marion and Sequeira, Gabriel and Misra, Diganta and Dhakal, Shreeya and Rystrøm, Jonathan and Solomatin, Roman and Çağatan, Ömer and Kundu, Akash and Bernstorff, Martin and Xiao, Shitao and Sukhlecha, Akshita and Pahwa, Bhavish and Poświata, Rafał and GV, Kranthi Kiran and Ashraf, Shawon and Auras, Daniel and Plüster, Björn and Harries, Jan Philipp and Magne, Loïc and Mohr, Isabelle and Hendriksen, Mariya and Zhu, Dawei and Gisserot-Boukhlef, Hippolyte and Aarsen, Tom and Kostkan, Jan and Wojtasik, Konrad and Lee, Taemin and Šuppa, Marek and Zhang, Crystina and Rocca, Roberta and Hamdy, Mohammed and Michail, Andrianos and Yang, John and Faysse, Manuel and Vatolin, Aleksei and Thakur, Nandan and Dey, Manan and Vasani, Dipam and Chitale, Pranjal and Tedeschi, Simone and Tai, Nguyen and Snegirev, Artem and Günther, Michael and Xia, Mengzhou and Shi, Weijia and Lù, Xing Han and Clive, Jordan and Krishnakumar, Gayatri and Maksimova, Anna and Wehrli, Silvan and Tikhonova, Maria and Panchal, Henil and Abramov, Aleksandr and Ostendorff, Malte and Liu, Zheng and Clematide, Simon and Miranda, Lester James and Fenogenova, Alena and Song, Guangyu and Safi, Ruqiya Bin and Li, Wen-Ding and Borghini, Alessia and Cassano, Federico and Su, Hongjin and Lin, Jimmy and Yen, Howard and Hansen, Lasse and Hooker, Sara and Xiao, Chenghao and Adlakha, Vaibhav and Weller, Orion and Reddy, Siva and Muennighoff, Niklas},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.13595v4},
  eprint = {2502.13595}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors