Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Yanzhao ZhangMingxin LiDingkun LongXin ZhangHuan LinBaosong YangPengjun XieAn YangDayiheng LiuJunyang Lin

article2025arXiv1,399 citations

Presents the open-source Qwen3 Embedding and reranking model suite across 0.6B to 8B parameters, setting new state-of-the-art performance on multilingual and code retrieval benchmarks through foundation-model-driven synthetic data generation and multi-stage training.

Listen

Modern natural language processing and information retrieval systems—such as enterprise search, recommendation engines, and retrieval-augmented generation—rely heavily on text embedding and reranking models. However, training these models to balance scalability, deep contextual comprehension, multilingual support, and alignment with specialized downstream tasks remains a substantial challenge. The article evaluates the Qwen3 Embedding series, an open-source suite of embedding and reranking models built on Qwen3 foundation models, to demonstrate how large foundation models can enhance representation learning across diverse operational requirements.

The authors implemented a multi-stage training pipeline leveraging the Qwen3-32B model to synthesize approximately 150 million diverse, weakly supervised text pairs across multiple languages, domains, and tasks. For the embedding models, this synthetic pre-training was followed by fine-tuning on a curated mix of roughly 7 million labeled pairs and 12 million high-quality synthetic pairs, culminating in a model merging technique using spherical linear interpolation across checkpoints. For reranking, the models were trained via supervised fine-tuning framed as binary relevance classification followed by model merging. The suite provides three parameter sizes (0.6B, 4B, and 8B) for both embedding and reranking, supporting flexible embedding dimensions and custom instruction awareness across sequence lengths up to 32,000 tokens.

Empirical evaluations demonstrate that the suite achieves state-of-the-art results across standard industry benchmarks. On the Massive Multilingual Text Embedding Benchmark (MTEB Multilingual), the flagship Qwen3-Embedding-8B achieved an overall task average of 70.58, surpassing proprietary alternatives such as Gemini-Embedding (68.37) and text-embedding-3-large (58.93). On the MTEB Code benchmark, Qwen3-Embedding-8B scored 80.68, outperforming top existing models by substantial margins. Even the compact Qwen3-Embedding-0.6B model proved highly competitive, achieving 64.33 on multilingual benchmarks and 75.41 on code retrieval, matching or exceeding much larger baseline models. Similarly, the Qwen3-Reranker models consistently improved ranking quality, with the 8B reranker delivering roughly 3 points of performance gain over the 0.6B variant on complex retrieval tasks.

These findings indicate that synthetic data generation and foundation-model pre-training can effectively substitute for costly, manual data curation while improving multilingual and cross-domain adaptability. Organizations can deploy smaller models like the 0.6B variant to minimize latency and compute expenses with minimal accuracy loss, or implement the 4B and 8B models for maximum retrieval precision in high-value enterprise applications. Ablation studies confirm that both the large-scale synthetic pre-training and checkpoint merging stages are critical drivers of this superior generalization.

Decision-makers can evaluate the open-source Apache 2.0-licensed models as drop-in enhancements for current retrieval and search architectures. Teams should select model tiers based on specific throughput and accuracy trade-offs, and conduct internal pilot evaluations on proprietary data distributions to validate real-world retrieval gains before full deployment.

  • Paper: On the Theoretical Limitations of Embedding-Based Retrieval, Orion Weller et al. (2026). This study analyzes the fundamental theoretical and dimensional limits of single-vector dense representations like those generated by Qwen3 Embedding.
  • Paper: Making Text Embedders Few-Shot Learners, Chaofan Li et al. (2025). This document explores extending LLM-based text embedders to become few-shot learners via in-context learning prompts without modifying the underlying model architecture.
  • Paper: Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini, Madhuri Shanbhogue et al. (2026). This work extends foundation-model-based embedding architectures from purely textual representations to unified native multimodal embeddings across text, images, video, and audio.
  • Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). This technical report details the multimodal vision-language progression of the Qwen3 ecosystem, extending Qwen3's underlying foundation capabilities into high-resolution spatial and video reasoning.
Cover for Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Abstract

In this work, we introduce the Qwen3 Embedding series, a significant advancement over its predecessor, the GTE-Qwen series, in text embedding and reranking capabilities, built upon the Qwen3 foundation models. Leveraging the Qwen3 LLMs' robust capabilities in multilingual text understanding and generation, our innovative multi-stage training pipeline combines large-scale unsupervised pre-training with supervised fine-tuning on high-quality datasets. Effective model merging strategies further ensure the robustness and adaptability of the Qwen3 Embedding series. During the training process, the Qwen3 LLMs serve not only as backbone models but also play a crucial role in synthesizing high-quality, rich, and diverse training data across multiple domains and languages, thus enhancing the training pipeline. The Qwen3 Embedding series offers a spectrum of model sizes (0.6B, 4B, 8B) for both embedding and reranking tasks, addressing diverse deployment scenarios where users can optimize for either efficiency or effectiveness. Empirical evaluations demonstrate that the Qwen3 Embedding series achieves state-of-the-art results across diverse benchmarks. Notably, it excels on the multilingual evaluation benchmark MTEB for text embedding, as well as in various retrieval tasks, including code retrieval, cross-lingual retrieval and multilingual retrieval. To facilitate reproducibility and promote community-driven research and development, the Qwen3 Embedding models are publicly available under the Apache 2.0 license.

Table of Contents

  • 1 Introduction
  • 2 Model Architecture
  • 3 Models Training
  • 3.1 Training Objective
  • 3.2 Multi-stage Training
  • 3.3 Synthetic Dataset
  • 4 Evaluation
  • 4.1 Settings
  • 4.2 Main Results
  • 4.3 Analysis
  • 5 Conclusion
  • References
  • A Appendix
  • A.1 Synthetic Data
  • A.2 Detail Results

Knowls

  1. Knowl 1 — Multi-Stage Training Pipeline for Qwen3 Embedding and Reranking Models

    model/method

    The training pipeline for the Qwen3 Embedding and Reranking series comprises distinct multi-stage processes designed to maximize retrieval accuracy, task generalization, and distribution robustness.

    For the Qwen3 text embedding models, the training follows a three-stage pipeline:

    1. Weakly Supervised Pre-Training: The model is trained on approximately 150 million pairs of multi-task synthetic text relevance data generated using Qwen3-32B across multiple domains, languages, and task types (retrieval, bitext mining, classification, semantic textual similarity) using an improved contrastive loss objective.
    2. Supervised Fine-Tuning (SFT): The pre-trained model is fine-tuned on a high-quality dataset combining approximately 7 million labeled pairs from public benchmarks (e.g., MS MARCO, NQ, HotpotQA, NLI, DuReader, T2T^2-Ranking, SimCLUE, MIRACL, MLDR, Mr.TyDi, Multi-CPR, CodeSearchNet) and approximately 12 million high-quality filtered synthetic pairs (selected via cosine similarity threshold >0.7> 0.7).
    3. Model Merging: Multiple intermediate checkpoints saved during the supervised fine-tuning stage are merged using spherical linear interpolation (SLERP) to mitigate task interference, balance data distributions, and improve out-of-distribution generalization.

    For the Qwen3 text reranking models, the training skips the weakly supervised pre-training phase and consists of a two-stage process: high-quality supervised fine-tuning directly on labeled and filtered task data using a point-wise binary classification cross-entropy loss, followed by SLERP-based model merging over intermediate fine-tuning checkpoints.

  2. Knowl 2 — InfoNCE-Based Contrastive Loss with In-Batch Negatives and False Negative Masking

    equation

    The Qwen3 text embedding models are trained using a modified InfoNCE contrastive loss over a batch of NN training instances. For an instance with query qiq_i, positive document di+d_i^+, and KK designated hard negatives di,k−d_{i,k}^-, the embedding loss is defined as:

    Lembedding=−1N∑i=1Nlog⁡exp⁡(s(qi,di+)/τ)Zi\mathcal{L}_{\text{embedding}} = -\frac{1}{N} \sum_{i=1}^N \log \frac{\exp(s(q_i, d_i^+)/\tau)}{Z_i}

    where s(u,v)=u⋅v∥u∥∥v∥s(u, v) = \frac{u \cdot v}{\|u\| \|v\|} denotes the cosine similarity function between vector representations, τ\tau is a temperature hyperparameter, and the partition function ZiZ_i aggregates similarity scores across the positive pair, hard negatives, and all cross-pair combinations in the batch:

    Zi=exp⁡(s(qi,di+)/τ)+∑k=1Kmikexp⁡(s(qi,di,k−)/τ)+∑j≠imijexp⁡(s(qi,qj)/τ)+∑j≠imijexp⁡(s(di+,dj)/τ)+∑j≠imijexp⁡(s(qi,dj)/τ)Z_i = \exp(s(q_i, d_i^+)/\tau) + \sum_{k=1}^K m_{ik} \exp(s(q_i, d_{i,k}^-)/\tau) + \sum_{j \neq i} m_{ij} \exp(s(q_i, q_j)/\tau) + \sum_{j \neq i} m_{ij} \exp(s(d_i^+, d_j)/\tau) + \sum_{j \neq i} m_{ij} \exp(s(q_i, d_j)/\tau)

    To prevent penalizing semantically valid alternate positives present in the batch, the mask factor mij∈{0,1}m_{ij} \in \{0, 1\} filters potential false negatives:

    mij={0if sij>s(qi,di+)+0.1ordj=di+1otherwisem_{ij} = \begin{cases} 0 & \text{if } s_{ij} > s(q_i, d_i^+) + 0.1 \quad \text{or} \quad d_j = d_i^+ \\ 1 & \text{otherwise} \end{cases}

    where sijs_{ij} represents the cosine similarity corresponding to (qi,dj)(q_i, d_j), (qi,qj)(q_i, q_j), (di+,dj)(d_i^+, d_j), or (qi,di,k−)(q_i, d_{i,k}^-).

  3. Knowl 3 — Architectural Designs and Prompt Formatting for Qwen3 Embedding and Reranking Models

    model/method

    The Qwen3 Embedding and Reranking models are built upon the dense autoregressive architectures of the Qwen3 foundation model family in three parameter configurations: 0.6B, 4B, and 8B parameters, supporting up to 32K token sequence context lengths.

    Embedding Models:

    • Feature causal attention with an appended [EOS] (<|endoftext|>) token at the end of the sequence. The dense representation is extracted directly from the hidden state of the final transformer layer at the [EOS] token position.
    • Queries and tasks incorporate task instructions formatted as {Instruction} {Query}<|endoftext|>. Documents are ingested directly without instruction prefixing to preserve standard corpus indexing efficiency.
    • Support Matryoshka Representation Learning (MRL), enabling truncation to lower vector dimensions at inference without retraining.

    Reranking Models:

    • Implement point-wise relevance classification within a single chat context formatted according to the standard Qwen chat template:
    <|im_start|>system
    Judge whether the Document meets the requirements based on the Query and the Instruct provided. Note that the answer can only be "yes" or "no".<|im_end|>
    <|im_start|>user
    <Instruct>: {Instruction}
    <Query>: {Query}
    <Document>: {Document}<|im_end|>
    <|im_start|>assistant
    <think>
    
    </think>
    
    
    
    • Relevance assessment evaluates the generation log-probabilities for the next token predicting "yes" vs. "no".
  4. Knowl 4 — Point-Wise Classification Objective and Scoring Formulation for Qwen3 Rerankers

    equation

    The Qwen3 text reranker evaluates query-document relevance by formulating reranking as a point-wise binary classification task conditioned on an explicit task instruction II, query qq, and candidate document dd.

    The final scalar relevance score score(q,d)∈(0,1)\text{score}(q, d) \in (0, 1) is computed from the unnormalized log-probabilities (or logits) assigned by the model to the generation tokens "yes" and "no":

    score(q,d)=exp⁡(P("yes"∣I,q,d))exp⁡(P("yes"∣I,q,d))+exp⁡(P("no"∣I,q,d))\text{score}(q, d) = \frac{\exp(P(\text{"yes"} \mid I, q, d))}{\exp(P(\text{"yes"} \mid I, q, d)) + \exp(P(\text{"no"} \mid I, q, d))}

    During training, the reranker is optimized using the Supervised Fine-Tuning (SFT) negative log-likelihood loss:

    Lreranking=−log⁡p(l∣I,q,d)\mathcal{L}_{\text{reranking}} = -\log p(l \mid I, q, d)

    where p(l∣I,q,d)p(l \mid I, q, d) is the model's predicted probability for the ground-truth label token l∈{"yes","no"}l \in \{\text{"yes"}, \text{"no"}\}, assigning higher likelihoods to "yes" for relevant pairs and "no" for irrelevant or hard negative pairs.

  5. Knowl 5 — Persona-Guided Multi-Dimensional Synthetic Relevance Data Generation

    model/method

    To generate large-scale multi-task training pairs for weak supervision and fine-tuning, a two-stage prompting pipeline is executed using the Qwen3-32B foundation model over a diverse multilingual web and text corpus.

    1. Configuration Stage: Given a raw source passage dd from the multilingual pre-training corpus, an auxiliary retriever selects the top 5 candidate user personas from Persona Hub that are most likely to interact with or seek information from dd. The foundation model selects the single most suitable character persona, a question type (keywords, acquire_knowledge, summary, yes_or_no, background), and a target difficulty level (high_school, university, phd), outputting a structured JSON configuration.

    2. Query Generation Stage: Conditioned on the selected persona, passage, question type, difficulty, target length constraint, and target output language, the LLM synthesizes a query simulating how that specific persona would retrieve the passage. This generates multilingual and cross-lingual text pairs across retrieval, bitext mining, semantic textual similarity, and classification tasks.

    3. Quality Filtering: From the initial pool of ~150 million generated pairs, cosine similarity filtering is applied to sampled subsets: pairs with cosine similarity score >0.7> 0.7 are extracted, yielding ~12 million high-quality synthetic instances for the supervised fine-tuning stage.

  6. Knowl 6 — Multilingual Benchmark Performance of Qwen3-Embedding on MMTEB

    data/table

    The Qwen3 text embedding models were evaluated on the Massive Multilingual Text Embedding Benchmark (MMTEB), spanning across 131 multilingual evaluation tasks in over 250 languages.

    Model Size Mean (Task) Mean (Type) Bitext Mining Classification Clustering Inst. Retrieval Multilabel Class. Pair Class. Rerank Retrieval STS
    NV-Embed-v2 7B 56.29 49.58 57.84 57.29 40.80 1.04 18.63 78.94 63.82 56.72 71.10
    GritLM-7B 7B 60.92 53.74 70.53 61.83 49.75 3.45 22.77 79.94 63.78 58.31 73.33
    BGE-M3 0.6B 59.56 52.18 79.11 60.35 40.88 -3.11 20.10 80.76 62.79 54.60 74.12
    multilingual-e5-large-instruct 0.6B 63.22 55.08 80.13 64.94 50.75 -0.40 22.91 80.86 62.61 57.12 76.81
    gte-Qwen2-7b-Instruct 7B 62.51 55.93 73.92 61.55 52.77 4.94 25.48 85.13 65.55 60.08 73.98
    text-embedding-3-large - 58.93 51.41 62.17 60.27 46.89 -2.68 22.03 79.17 63.89 59.27 71.68
    Cohere-embed-multilingual-v3.0 - 61.12 53.23 70.50 62.95 46.89 -1.89 22.74 79.88 64.07 59.16 74.80
    Gemini Embedding - 68.37 59.59 79.28 71.82 54.59 5.18 29.16 83.63 65.58 67.71 79.40
    Qwen3-Embedding-0.6B 0.6B 64.33 56.00 72.22 66.83 52.33 5.09 24.59 80.83 61.41 64.64 76.17
    Qwen3-Embedding-4B 4B 69.45 60.86 79.36 72.33 57.15 11.56 26.77 85.05 65.08 69.60 80.86
    Qwen3-Embedding-8B 8B 70.58 61.69 80.89 74.00 57.65 10.06 28.66 86.40 65.63 70.88 81.08

    Qwen3-Embedding-8B achieved the top overall score on MMTEB (70.58 Mean Task, 61.69 Mean Type), surpassing proprietary models such as Gemini Embedding (68.37) and open-source models like GritLM-7B (60.92). The 0.6B model achieved 64.33 Mean Task, outperforming larger baseline models such as multilingual-e5-large-instruct (63.22) and gte-Qwen2-7b-Instruct (62.51).

  7. Knowl 7 — Text Reranking Benchmark Evaluation Across Retrieval Domains

    data/table

    The Qwen3-Reranker series was evaluated on candidate pools formed by the top-100 results retrieved by Qwen3-Embedding-0.6B across English retrieval (MTEB-R), Chinese retrieval (CMTEB-R), Multilingual retrieval (MMTEB-R, MLDR), Code retrieval (MTEB-Code), and complex instruction following (FollowIR).

    Model Parameters MTEB-R CMTEB-R MMTEB-R MLDR MTEB-Code FollowIR
    Qwen3-Embedding-0.6B (First-stage retrieval) 0.6B 61.82 71.02 64.64 50.26 75.41 5.09
    Jina-multilingual-reranker-v2-base 0.3B 58.22 63.37 63.73 39.66 58.98 -0.68
    gte-multilingual-reranker-base 0.3B 59.51 74.08 59.44 66.33 54.18 -1.64
    BGE-reranker-v2-m3 0.6B 57.03 72.16 58.36 59.51 41.38 -0.01
    Qwen3-Reranker-0.6B 0.6B 65.80 71.31 66.36 67.28 73.42 5.41
    Qwen3-Reranker-4B 4B 69.76 75.94 72.74 69.97 81.20 14.84
    Qwen3-Reranker-8B 8B 69.02 77.45 72.94 70.19 81.22 8.05

    Qwen3-Reranker-8B achieved the highest scores on CMTEB-R (77.45), MMTEB-R (72.94), MLDR (70.19), and MTEB-Code (81.22). Qwen3-Reranker-4B achieved the highest score on complex instruction following on FollowIR (14.84), while significantly outperforming previous baseline rerankers like BGE-reranker-v2-m3 (-0.01) and gte-multilingual-reranker-base (-1.64).

  8. Knowl 8 — Performance on MTEB English v2, Chinese (CMTEB), and Code Benchmarks

    data/table

    The Qwen3 Embedding models were evaluated on English (MTEB v2, 41 tasks), Chinese (CMTEB, 32 tasks), and Code (MTEB-Code, 12 retrieval tasks reporting nDCG@10):

    Model Size Embedding Dim MTEB (Eng, v2) CMTEB MTEB (Code)
    Mean (Task) Mean (Type) Mean (Task) Mean (Type) nDCG@10
    NV-Embed-v2 7B 4096 69.81 65.00 63.00 62.00 63.74
    GritLM-7B 7B 4096 67.07 63.22 - - 73.60
    multilingual-e5-large-instruct 0.6B 1024 65.53 61.21 58.08 58.24 65.00
    gte-Qwen2-7b-instruct 7B 3584 70.72 65.77 71.62 72.19 62.17
    text-embedding-3-large - 3072 66.43 62.15 - - 58.95
    Gemini Embedding - 3072 73.30 67.67 - - 74.66
    Qwen3-Embedding-0.6B 0.6B 1024 70.70 64.88 66.33 67.44 75.41
    Qwen3-Embedding-4B 4B 2560 74.60 68.09 72.26 73.50 80.06
    Qwen3-Embedding-8B 8B 4096 75.22 68.70 73.84 75.00 80.68

    Qwen3-Embedding-8B established new state-of-the-art results across English (75.22 Mean Task), Chinese (73.84 Mean Task), and Code (80.68 nDCG@10), outperforming Gemini Embedding (73.30 on English, 74.66 on Code). Qwen3-Embedding-0.6B reached 75.41 on MTEB Code, surpassing all evaluated commercial and open-source 7B-parameter baselines.

  9. Knowl 9 — Ablation Study on Synthetic Weak Supervision and Model Merging

    empirical result

    An ablation study conducted on Qwen3-Embedding-0.6B demonstrates the individual contributions of large-scale weakly supervised pre-training on synthetic data and post-SFT model merging across four evaluation benchmarks:

    1. Pre-Training on Synthetic Data Only: Training solely on the ~150M weakly supervised synthetic data pairs (without subsequent supervised fine-tuning) yields strong baseline performance: MMTEB: 58.49, MTEB (Eng, v2): 60.63, CMTEB: 59.78, MTEB (Code, v1): 66.79.
    2. Removal of Synthetic Pre-Training: Removing Stage 1 synthetic data training and training only on supervised datasets causes performance to drop across all benchmarks: MMTEB drops from 64.33 to 61.21 (-3.12), MTEB (Eng, v2) drops from 70.70 to 65.59 (-5.11), CMTEB drops from 66.33 to 63.37 (-2.96), and MTEB Code drops from 75.41 to 74.58 (-0.83).
    3. Removal of Model Merging: Training the full pipeline but replacing SLERP model merging with standard task-balanced data sampling in fine-tuning degrades performance: MMTEB drops to 62.56 (-1.77), MTEB (Eng, v2) drops to 68.18 (-2.52), CMTEB drops to 64.76 (-1.57), and MTEB Code drops to 74.89 (-0.52).

    The results confirm that synthetic pre-training is essential for broad generalization and SLERP model merging is necessary to resolve multi-task distribution conflicts.

  10. Knowl 10 — Structural Specifications of the Qwen3 Embedding Series

    definition

    The Qwen3 Embedding and Reranker series provides models across three scale tiers (0.6B, 4B, and 8B parameters) with the following architectural specifications:

    • Qwen3-Embedding-0.6B / Qwen3-Reranker-0.6B: 0.6 billion parameters, 28 transformer layers, maximum context sequence length of 32,768 (32K) tokens. Embedding dimension is 1024 with Matryoshka Representation Learning (MRL) dimension reduction support.
    • Qwen3-Embedding-4B / Qwen3-Reranker-4B: 4.0 billion parameters, 36 transformer layers, maximum context sequence length of 32K tokens. Embedding dimension is 2560 with MRL dimension reduction support.
    • Qwen3-Embedding-8B / Qwen3-Reranker-8B: 8.0 billion parameters, 36 transformer layers, maximum context sequence length of 32K tokens. Embedding dimension is 4096 with MRL dimension reduction support.

    All embedding and reranking variants natively support custom task instructions (instruction-awareness) to adapt similarity criteria dynamically at inference time.

Coverage note — No substantial contributed material was omitted. All primary training methodologies, mathematical loss formulas, architectural templates, synthetic data generation pipelines, ablation analyses, and benchmark evaluations across multilingual, English, Chinese, code, and reranking tasks are covered.

References

  1. 1.Jianlyu Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. In Findings of the Association for Computational Linguistics: ACL 2024, pp. 2318–2335, Bangkok, Thailand, August 2024. Association for Computational Linguistics. URL https://aclanthology.org/2024.findings-acl.137/.
  2. 2.Kenneth Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzemiński, Genta Indra Winata, et al. MMTEB: Massive multilingual text embedding benchmark. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=zl3pfz4VCV.
  3. 3.Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu. Scaling synthetic data creation with 1,000,000,000 personas. arXiv preprint arXiv:2406.20094, 2024.
  4. 4.Jui-Ting Huang, Ashish Sharma, Shuying Sun, Li Xia, David Zhang, Philip Pronin, Janani Padmanabhan, Giuseppe Ottaviano, and Linjun Yang. Embedding-based retrieval in facebook search. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 2553–2561, 2020.
  5. 5.Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024.
  6. 6.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick SH Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. Dense passage retrieval for open-domain question answering. In EMNLP (1), pp. 6769–6781, 2020.
  7. 7.Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. Nv-embed: Improved techniques for training llms as generalist embedding models. arXiv preprint arXiv:2405.17428, 2024.
  8. 8.Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. NV-embed: Improved techniques for training LLMs as generalist embedding models. In The Thirteenth International Conference on Learning Representations, 2025a. URL https://openreview.net/forum?id=lgsyLSsDRe.
  9. 9.Jinhyuk Lee, Feiyang Chen, Sahil Dua, Daniel Cer, Madhuri Shanbhogue, Iftekhar Naim, Gustavo Hernández Ábrego, Zhe Li, Kaifeng Chen, Henrique Schechter Vera, et al. Gemini embedding: Generalizable embeddings from gemini. arXiv preprint arXiv:2503.07891, 2025b.
  10. 10.Mingxin Li, Zhijie Nie, Yanzhao Zhang, Dingkun Long, Richong Zhang, and Pengjun Xie. Improving general text embedding model: Tackling task conflict and data imbalance through model merging. arXiv preprint arXiv:2410.15035, 2024.
  11. 11.Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. Towards general text embeddings with multi-stage contrastive learning, 2023. URL https://arxiv.org/abs/2308.03281.
  12. 12.Xueguang Ma, Xinyu Zhang, Ronak Pradeep, and Jimmy Lin. Zero-shot listwise document reranking with a large language model. arXiv preprint arXiv:2305.02156, 2023.
  13. 13.Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. MTEB: Massive text embedding benchmark. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pp. 2014–2037, Dubrovnik, Croatia, May 2023. Association for Computational Linguistics. URL https://aclanthology.org/2023.eacl-main.148/.
  14. 14.Niklas Muennighoff, Hongjin SU, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela. Generative representational instruction tuning. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=BC4lIvfSzv.
  15. 15.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  16. 16.Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. Rankvicuna: Zero-shot listwise document reranking with open-source large language models. arXiv preprint arXiv:2309.15088, 2023.
  17. 17.Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 3982–3992, Hong Kong, China, November 2019. Association for Computational Linguistics. URL https://aclanthology.org/D19-1410/.
  18. 18.Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A Smith, Luke Zettlemoyer, and Tao Yu. One embedder, any task: Instruction-finetuned text embeddings. In Findings of the Association for Computational Linguistics: ACL 2023, pp. 1102–1121, 2023.
  19. 19.Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. Text embeddings by weakly-supervised contrastive pre-training, 2022. URL https://arxiv.org/abs/2212.03533.
  20. 20.Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. Improving text embeddings with large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 11897–11916, Bangkok, Thailand, August 2024. Association for Computational Linguistics. URL https://aclanthology.org/2024.acl-long.642/.
  21. 21.Orion Weller, Benjamin Chang, Sean MacAvaney, Kyle Lo, Arman Cohan, Benjamin Van Durme, Dawn Lawrie, and Luca Soldaini. Followir: Evaluating and teaching information retrieval models to follow instructions. arXiv preprint arXiv:2403.15246, 2024.
  22. 22.Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. C-pack: Packed resources for general chinese embeddings. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’24, pp. 641–649, New York, NY, USA, 2024. Association for Computing Machinery. URL https://doi.org/10.1145/3626772.3657878.
  23. 23.An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025.
  24. 24.Longhui Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, and Min Zhang. A two-stage adaptation of large language models for text ranking. In Findings of the Association for Computational Linguistics ACL 2024, pp. 11880–11891, 2024a.
  25. 25.Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li, and Min Zhang. mGTE: Generalized long-context text representation and reranking models for multilingual text retrieval. In Franck Dernoncourt, Daniel Preoţiuc-Pietro, and Anastasia Shimorina (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, pp. 1393–1412, Miami, Florida, US, November 2024b. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-industry.103. URL https://aclanthology.org/2024.emnlp-industry.103/.
  26. 26.Wayne Xin Zhao, Jing Liu, Ruiyang Ren, and Ji-Rong Wen. Dense text retrieval based on pretrained language models: A survey. ACM Transactions on Information Systems, 42(4):1–60, 2024.
  27. 27.Xiangyu Zhao, Maolin Wang, Xinjian Zhao, Jiansheng Li, Shucheng Zhou, Dawei Yin, Qing Li, Jiliang Tang, and Ruocheng Guo. Embedding in recommender systems: A survey. arXiv preprint arXiv:2310.18608, 2023.
  28. 28.Shengyao Zhuang, Honglei Zhuang, Bevan Koopman, and Guido Zuccon. A setwise approach for effective and highly efficient zero-shot ranking with large language models. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 38–47, 2024.

Citation

MLA
Zhang, Y., et al. “Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models”. arXiv, 2025, http://arxiv.org/abs/2506.05176v3.
APA
Zhang, Y., Li, M., Long, D., Zhang, X., Lin, H., Yang, B., Xie, P., Yang, A., Liu, D., Lin, J., Huang, F., & Zhou, J. (2025). Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models. arXiv. http://arxiv.org/abs/2506.05176v3
Chicago
Zhang, Y., M. Li, D. Long, et al. 2025. “Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models”. arXiv. http://arxiv.org/abs/2506.05176v3.
Harvard
Zhang, Y. et al. (2025) “Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2506.05176v3.
Vancouver
1. Zhang Y, Li M, Long D, et al (2025) Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models. arXiv

BibTeX

@article{zhang2025qwen3,
  title = {Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models},
  author = {Zhang, Yanzhao and Li, Mingxin and Long, Dingkun and Zhang, Xin and Lin, Huan and Yang, Baosong and Xie, Pengjun and Yang, An and Liu, Dayiheng and Lin, Junyang and Huang, Fei and Zhou, Jingren},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2506.05176v3},
  eprint = {2506.05176}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors