Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Yanzhao ZhangMingxin LiDingkun LongXin ZhangHuan LinBaosong YangPengjun XieAn YangDayiheng LiuJunyang Lin
Presents the open-source Qwen3 Embedding and reranking model suite across 0.6B to 8B parameters, setting new state-of-the-art performance on multilingual and code retrieval benchmarks through foundation-model-driven synthetic data generation and multi-stage training.
Modern natural language processing and information retrieval systems—such as enterprise search, recommendation engines, and retrieval-augmented generation—rely heavily on text embedding and reranking models. However, training these models to balance scalability, deep contextual comprehension, multilingual support, and alignment with specialized downstream tasks remains a substantial challenge. The article evaluates the Qwen3 Embedding series, an open-source suite of embedding and reranking models built on Qwen3 foundation models, to demonstrate how large foundation models can enhance representation learning across diverse operational requirements.
The authors implemented a multi-stage training pipeline leveraging the Qwen3-32B model to synthesize approximately 150 million diverse, weakly supervised text pairs across multiple languages, domains, and tasks. For the embedding models, this synthetic pre-training was followed by fine-tuning on a curated mix of roughly 7 million labeled pairs and 12 million high-quality synthetic pairs, culminating in a model merging technique using spherical linear interpolation across checkpoints. For reranking, the models were trained via supervised fine-tuning framed as binary relevance classification followed by model merging. The suite provides three parameter sizes (0.6B, 4B, and 8B) for both embedding and reranking, supporting flexible embedding dimensions and custom instruction awareness across sequence lengths up to 32,000 tokens.
Empirical evaluations demonstrate that the suite achieves state-of-the-art results across standard industry benchmarks. On the Massive Multilingual Text Embedding Benchmark (MTEB Multilingual), the flagship Qwen3-Embedding-8B achieved an overall task average of 70.58, surpassing proprietary alternatives such as Gemini-Embedding (68.37) and text-embedding-3-large (58.93). On the MTEB Code benchmark, Qwen3-Embedding-8B scored 80.68, outperforming top existing models by substantial margins. Even the compact Qwen3-Embedding-0.6B model proved highly competitive, achieving 64.33 on multilingual benchmarks and 75.41 on code retrieval, matching or exceeding much larger baseline models. Similarly, the Qwen3-Reranker models consistently improved ranking quality, with the 8B reranker delivering roughly 3 points of performance gain over the 0.6B variant on complex retrieval tasks.
These findings indicate that synthetic data generation and foundation-model pre-training can effectively substitute for costly, manual data curation while improving multilingual and cross-domain adaptability. Organizations can deploy smaller models like the 0.6B variant to minimize latency and compute expenses with minimal accuracy loss, or implement the 4B and 8B models for maximum retrieval precision in high-value enterprise applications. Ablation studies confirm that both the large-scale synthetic pre-training and checkpoint merging stages are critical drivers of this superior generalization.
Decision-makers can evaluate the open-source Apache 2.0-licensed models as drop-in enhancements for current retrieval and search architectures. Teams should select model tiers based on specific throughput and accuracy trade-offs, and conduct internal pilot evaluations on proprietary data distributions to validate real-world retrieval gains before full deployment.
- Paper: Qwen3 Technical Report, An Yang et al. (2025). This report details the architectural and training innovations of the Qwen3 foundation models that directly serve as the backbone and synthetic data engine for Qwen3 Embedding.
- Paper: M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation, Jianlv Chen et al. (2024). Reading M3-Embedding provides foundational context on multi-stage pre-training, multilingual scaling, and self-distillation techniques that modern LLM-based text embedders build upon.
- Paper: NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models, Chankyu Lee et al. (2025). This paper establishes key techniques for converting decoder-only large language models into general-purpose text embedding systems via multi-stage contrastive instruction tuning.
- Paper: Text Embeddings by Weakly-Supervised Contrastive Pre-training, Liang Wang et al. (2022). This work introduces weakly supervised contrastive pre-training on large-scale text pairs, establishing the multi-stage training paradigm central to the Qwen embedding pipeline.
- Paper: Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval, Lee Xiong et al. (2021). This paper details hard negative mining and contrastive learning mechanisms that are critical to optimizing dense retrieval models before fine-tuning.
- Paper: Qwen2.5 Technical Report, Qwen et al. (2024). This report covers the preceding generation of Qwen foundation models and their architectural optimizations, establishing the immediate lineage for the Qwen3 family.
- Paper: Qwen2 Technical Report, An Yang et al. (2024). This technical report outlines the foundational pre-training and alignment strategies of the Qwen series that inform subsequent Qwen foundation and embedding architectures.
- Paper: Making Monolingual Sentence Embeddings Multilingual Using Knowledge Distillation, Nils Reimers et al. (2020). This work explains how cross-lingual knowledge distillation enables strong multilingual embedding capabilities across diverse low-resource languages.
- Paper: On the Theoretical Limitations of Embedding-Based Retrieval, Orion Weller et al. (2026). This study analyzes the fundamental theoretical and dimensional limits of single-vector dense representations like those generated by Qwen3 Embedding.
- Paper: Making Text Embedders Few-Shot Learners, Chaofan Li et al. (2025). This document explores extending LLM-based text embedders to become few-shot learners via in-context learning prompts without modifying the underlying model architecture.
- Paper: Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini, Madhuri Shanbhogue et al. (2026). This work extends foundation-model-based embedding architectures from purely textual representations to unified native multimodal embeddings across text, images, video, and audio.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). This technical report details the multimodal vision-language progression of the Qwen3 ecosystem, extending Qwen3's underlying foundation capabilities into high-resolution spatial and video reasoning.
