Qwen2 Technical Report
An YangBaosong YangBinyuan HuiBo ZhengBowen YuChang ZhouChengpeng LiChengyuan LiDayiheng LiuFei Huang
Presents the Qwen2 family of open-weight dense and Mixture-of-Experts large language models ranging from 0.5B to 72B parameters, detailing training strategies and architectures that achieve performance competitive with leading proprietary systems across coding, mathematics, reasoning, and multilingual benchmarks in approximately 30 languages.
The rapid advancement of artificial intelligence has intensified demand for accessible, high-performing foundation models that can compete with leading proprietary systems. While commercial closed-source models frequently set industry benchmarks, open-weight alternatives often trail behind in specialized reasoning, multilingual fluency, and extended context handling. Organizations seeking deployable, cost-efficient artificial intelligence solutions require robust open models that span multiple computing footprints—from mobile edge hardware to enterprise GPU clusters.
To address this need, the article presents the design, training, and evaluation of Qwen2, an open-weight family of large language models. The primary objective is to demonstrate that scalable pre-training, architectural optimizations, and automated post-training alignment can elevate open models to performance levels that match or surpass leading proprietary and open-source alternatives.
The development process evaluated models across five specific configurations: four standard dense models sized at 0.5 billion, 1.5 billion, 7 billion, and 72 billion parameters, alongside a 57-billion-parameter Mixture-of-Experts model that dynamically activates 14 billion parameters per token. The foundational models were pre-trained on a high-quality dataset containing over 7 trillion tokens covering approximately 30 languages, with enhanced representation of programming and mathematical corpora. To extend context processing up to 131,072 tokens, the approach integrated Grouped Query Attention, Dual Chunk Attention, and attention-rescaling mechanisms. Post-training leveraged over 500,000 instruction examples through supervised fine-tuning, automated synthetic feedback, and reinforcement learning.
Evaluations across standard benchmarks revealed several major findings. First, the flagship 72-billion-parameter base model outperformed prominent open baselines, including Llama-3-70B, scoring 84.2 on general language understanding (MMLU) and 64.6 on coding (HumanEval). Second, the instruction-tuned 72-billion variant achieved state-of-the-art alignment results, recording a 9.12 on MT-Bench and 48.1 on Arena-Hard, while demonstrating safety rejection rates superior to GPT-4 on high-risk prompts. Third, the 57-billion Mixture-of-Experts model matched the general performance of dense 30-billion-parameter models while reducing computational load by activating only 14 billion parameters per token. Fourth, smaller models demonstrated strong efficiency; the 1.5-billion model consistently surpassed comparable small baselines, confirming that data scale enhances sub-billion and low-billion parameter architectures.
These findings indicate that high-performing artificial intelligence capabilities can be achieved without relying solely on massive, proprietary architectures. For enterprise decision-makers, this translates into lower inference costs, reduced memory footprints via attention optimizations, and flexible deployment options ranging from mobile devices to cloud infrastructure. The model family offers immediate utility for multilingual applications, long-document processing, and automated coding without sacrificing safety compliance.
Organizations should consider deploying Qwen2 models according to specific operational needs: the 0.5-billion and 1.5-billion models for on-device and edge applications, the Mixture-of-Experts or 7-billion models for balanced cost-to-throughput enterprise tasks, and the 72-billion model for heavy reasoning and coding workloads. Further post-training refinement is recommended for smaller variants, particularly the 7-billion model, to close remaining gaps in complex instruction following.
The findings are bounded by certain limitations, including a minor performance lag behind leading proprietary models on specific English comprehension tasks and the standard challenges of automated safety filtering for nuanced content. However, rigorous data decontamination analyses showed consistent performance between original and uncontaminated test sets, supporting high confidence in the reported benchmark capabilities.
- Paper: Qwen Technical Report, Jinze Bai et al. (2023). It introduces the initial foundational architecture, pre-training corpus, and instruction-tuning framework of the Qwen series upon which Qwen2 directly iterates.
- Paper: Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, Jinze Bai et al. (2023). It establishes the core architectural design and multimodal training pipeline that underpins the multimodal extensions referenced in the Qwen2 series.
- Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models, Hugo Touvron et al. (2023). It outlines key architectural standards and reinforcement learning alignment practices widely adopted and benchmarked against by Qwen2.
- Paper: DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models, Damai Dai et al. (2024). It details fine-grained expert routing and shared-expert design principles essential to understanding the Mixture-of-Experts architecture utilized in Qwen2-MoE.
- Paper: AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration, Ji Lin et al. (2024). It introduces activation-aware weight quantization, providing the technical basis for the low-bit quantization and efficient deployment methods applied to Qwen2.
- Paper: Qwen2.5 Technical Report, Qwen et al. (2024). It directly extends the Qwen2 model family by scaling context lengths up to 1 million tokens and advancing post-training reasoning capabilities.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). It adapts the Qwen2 base language models into multimodal vision-language systems by incorporating dynamic resolution perception and 3D rotary embeddings.
- Paper: Qwen3 Technical Report, An Yang et al. (2025). It represents the next major generational leap of the model family, scaling up sparse MoE architectures and integrating native long Chain-of-Thought reasoning.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). It continues the Qwen evolution by building advanced multimodal reasoning and 256K interleaved context support on top of subsequent Qwen language backbones.
- Paper: VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models, Sen Xu et al. (2026). It explores post-training verifiable reasoning by applying curriculum fine-tuning and reinforcement learning to a compact Qwen2-derived base model.
- Paper: LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models, Yaowei Zheng et al. (2024). It provides a unified fine-tuning and adaptation framework that practically operationalizes open-weight foundation models like Qwen2 across downstream tasks.
