Qwen2 Technical Report

An YangBaosong YangBinyuan HuiBo ZhengBowen YuChang ZhouChengpeng LiChengyuan LiDayiheng LiuFei Huang

article2024arXiv2,474 citations

Presents the Qwen2 family of open-weight dense and Mixture-of-Experts large language models ranging from 0.5B to 72B parameters, detailing training strategies and architectures that achieve performance competitive with leading proprietary systems across coding, mathematics, reasoning, and multilingual benchmarks in approximately 30 languages.

Listen

The rapid advancement of artificial intelligence has intensified demand for accessible, high-performing foundation models that can compete with leading proprietary systems. While commercial closed-source models frequently set industry benchmarks, open-weight alternatives often trail behind in specialized reasoning, multilingual fluency, and extended context handling. Organizations seeking deployable, cost-efficient artificial intelligence solutions require robust open models that span multiple computing footprints—from mobile edge hardware to enterprise GPU clusters.

To address this need, the article presents the design, training, and evaluation of Qwen2, an open-weight family of large language models. The primary objective is to demonstrate that scalable pre-training, architectural optimizations, and automated post-training alignment can elevate open models to performance levels that match or surpass leading proprietary and open-source alternatives.

The development process evaluated models across five specific configurations: four standard dense models sized at 0.5 billion, 1.5 billion, 7 billion, and 72 billion parameters, alongside a 57-billion-parameter Mixture-of-Experts model that dynamically activates 14 billion parameters per token. The foundational models were pre-trained on a high-quality dataset containing over 7 trillion tokens covering approximately 30 languages, with enhanced representation of programming and mathematical corpora. To extend context processing up to 131,072 tokens, the approach integrated Grouped Query Attention, Dual Chunk Attention, and attention-rescaling mechanisms. Post-training leveraged over 500,000 instruction examples through supervised fine-tuning, automated synthetic feedback, and reinforcement learning.

Evaluations across standard benchmarks revealed several major findings. First, the flagship 72-billion-parameter base model outperformed prominent open baselines, including Llama-3-70B, scoring 84.2 on general language understanding (MMLU) and 64.6 on coding (HumanEval). Second, the instruction-tuned 72-billion variant achieved state-of-the-art alignment results, recording a 9.12 on MT-Bench and 48.1 on Arena-Hard, while demonstrating safety rejection rates superior to GPT-4 on high-risk prompts. Third, the 57-billion Mixture-of-Experts model matched the general performance of dense 30-billion-parameter models while reducing computational load by activating only 14 billion parameters per token. Fourth, smaller models demonstrated strong efficiency; the 1.5-billion model consistently surpassed comparable small baselines, confirming that data scale enhances sub-billion and low-billion parameter architectures.

These findings indicate that high-performing artificial intelligence capabilities can be achieved without relying solely on massive, proprietary architectures. For enterprise decision-makers, this translates into lower inference costs, reduced memory footprints via attention optimizations, and flexible deployment options ranging from mobile devices to cloud infrastructure. The model family offers immediate utility for multilingual applications, long-document processing, and automated coding without sacrificing safety compliance.

Organizations should consider deploying Qwen2 models according to specific operational needs: the 0.5-billion and 1.5-billion models for on-device and edge applications, the Mixture-of-Experts or 7-billion models for balanced cost-to-throughput enterprise tasks, and the 72-billion model for heavy reasoning and coding workloads. Further post-training refinement is recommended for smaller variants, particularly the 7-billion model, to close remaining gaps in complex instruction following.

The findings are bounded by certain limitations, including a minor performance lag behind leading proprietary models on specific English comprehension tasks and the standard challenges of automated safety filtering for nuanced content. However, rigorous data decontamination analyses showed consistent performance between original and uncontaminated test sets, supporting high confidence in the reported benchmark capabilities.

Cover for Qwen2 Technical Report

Abstract

This report introduces the Qwen2 series, the latest addition to our large language models and large multimodal models. We release a comprehensive suite of foundational and instruction-tuned language models, encompassing a parameter range from 0.5 to 72 billion, featuring dense models and a Mixture-of-Experts model. Qwen2 surpasses most prior open-weight models, including its predecessor Qwen1.5, and exhibits competitive performance relative to proprietary models across diverse benchmarks on language understanding, generation, multilingual proficiency, coding, mathematics, and reasoning.

The flagship model, Qwen2-72B, showcases remarkable performance: 84.2 on MMLU, 37.9 on GPQA, 64.6 on HumanEval, 89.5 on GSM8K, and 82.4 on BBH as a base language model. The instruction-tuned variant, Qwen2-72B-Instruct, attains 9.1 on MT-Bench, 48.1 on Arena-Hard, and 35.7 on LiveCodeBench. Moreover, Qwen2 demonstrates robust multilingual capabilities, proficient in approximately 30 languages, spanning English, Chinese, Spanish, French, German, Arabic, Russian, Korean, Japanese, Thai, Vietnamese, and more, underscoring its versatility and global reach.

To foster community innovation and accessibility, we have made the Qwen2 model weights openly available on Hugging Face and ModelScope, and the supplementary materials including example code on GitHub. These platforms also include resources for quantization, fine-tuning, and deployment, facilitating a wide range of applications and research endeavors.

Table of Contents

  • 1 Introduction
  • 2 Tokenizer & Model
  • 2.1 Tokenizer
  • 2.2 Model Architecture
  • 2.2.1 Qwen2 Dense Model
  • 2.2.2 Qwen2 Mixture-of-experts Model
  • 2.2.3 Model Configuration
  • 3 Pre-training
  • 3.1 Pre-training Data
  • 3.2 Long-context Training
  • 4 Post-training
  • 4.1 Post-training Data
  • 4.1.1 Collaborative Data Annotation
  • 4.1.2 Automated Data Synthesis
  • 4.2 Supervised Fine-tuning
  • 4.3 Reinforcement Learning from Human Feedback
  • 5 Evaluation
  • 5.1 Base Language Models
  • 5.1.1 Core Capabilities
  • 5.2 Instruction-tuned Model
  • 5.2.1 Open Benchmark Evaluation
  • 5.2.2 In-house Automatic Evaluation
  • 5.2.3 Long Context Capabilities
  • 5.2.4 Multilingual Evaluation
  • 5.2.5 Safety & Responsibility
  • 5.2.6 Contamination Analysis
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Qwen2 Architecture and Model Configurations

    data/table

    The Qwen2 family comprises four dense autoregressive language models (0.5B, 1.5B, 7B, and 72B) and one Mixture-of-Experts model (Qwen2-57B-A14B). All models utilize the Transformer architecture with causal self-attention, SwiGLU activations, Rotary Position Embeddings (RoPE), QKV bias in attention projections, Pre-RMSNorm normalization, and Grouped Query Attention (GQA) to optimize Key-Value (KV) cache memory during inference. The vocabulary uses byte-level byte-pair encoding (BPE) with 151,643 regular tokens and 3 control tokens (effective vocabulary size is 151,646).

    Configuration 0.5B 1.5B 7B 72B 57B-A14B
    Hidden Size 896 1,536 3,584 8,192 3,584
    # Layers 24 28 28 80 28
    # Query Heads 14 12 28 64 28
    # KV Heads 2 2 4 8 4
    Head Size 64 128 128 128 128
    Intermediate Size 4,864 8,960 18,944 29,568 2,560
    # Routed Experts - - - - 64
    # Activated Experts - - - - 8
    # Shared Experts - - - - 8
    Embedding Tying True True False False False
    Vocabulary Size 151,646 151,646 151,646 151,646 151,646
    # Pre-trained Tokens 12T 7T 7T 7T 4.5T

    The intermediate size for Qwen2-57B-A14B refers to the hidden dimension of each individual expert. The 8 activated experts count excludes the 8 shared experts that are always active for every token.

  2. Knowl 2 — Qwen2-57B-A14B Mixture-of-Experts Architecture and Initialization

    model/method

    The Qwen2-57B-A14B Mixture-of-Experts (MoE) architecture replaces the standard feed-forward network (FFN) with a collection of fine-grained routed experts and shared experts. For an input representation xx, a gating network G(x)G(x) assigns routing probabilities over nn experts:

    p=softmax(G(x))p = \text{softmax}(G(x))

    y=∑i∈topk(p)piEi(x)y = \sum_{i \in \text{top}_k(p)} p_i E_i(x)

    where Ei(x)E_i(x) denotes the output of the ii-th expert FFN and topk(p)\text{top}_k(p) selects the indices of the kk highest routing probabilities. Qwen2-57B-A14B employs n=64n=64 routed experts (k=8k=8 activated per token) alongside 8 shared experts that are unconditionally executed for each token.

    To initialize the MoE model from a dense base model (Qwen2-7B), an upcycling procedure is applied:

    1. Given the target expert intermediate dimension hEh_E, the total number of experts nn, and the original dense model FFN intermediate dimension hFFNh_{\text{FFN}}, the original FFN layer is duplicated ⌈n×hE/hFFN⌉\lceil n \times h_E / h_{\text{FFN}} \rceil times.
    2. To induce diversity across expert copies, parameters are shuffled along the intermediate dimension within each duplicated FFN.
    3. The individual fine-grained expert weights of dimension hEh_E are sliced from the shuffled copies, and any excess dimensions are discarded.
    4. For each fine-grained expert, 50%50\% of its parameters are randomly reinitialized to encourage representational exploration during subsequent pre-training.
  3. Knowl 3 — Pre-Training Data and Long-Context Extension Pipeline

    model/method

    The pre-training dataset for Qwen2 comprises over 7 trillion tokens covering approximately 30 languages (including English, Chinese, Spanish, French, German, Arabic, Russian, Korean, Japanese, Thai, and Vietnamese), with enriched proportions of code and mathematics data. The 0.5B model is trained on 12 trillion tokens, while the 57B-A14B MoE model receives 4.5 trillion additional tokens after upcycling from the 7B dense model.

    To extend the native context window:

    1. Two-Stage Context Pre-training: During the final phase of pre-training, the context window is expanded from 4,096 tokens to 32,768 tokens using lengthy, high-quality documents.
    2. RoPE Base Frequency Modification: The base frequency θ\theta of the Rotary Position Embedding (RoPE) is increased from 10,00010{,}000 to 1,000,0001{,}000{,}000.
    3. Inference Length Extrapolation: For sequence lengths beyond 32,768 tokens (up to 131,072 tokens), YaRN (Yet another RoPE extensioN) is applied to rescale attention weights, and Dual Chunk Attention (DCA) is used to partition sequences into manageable chunks while preserving relative intra-chunk and inter-chunk position representations.
  4. Knowl 4 — Automated Post-Training Alignment Data Construction

    model/method

    Qwen2 post-training minimizes manual human labeling by pairing collaborative data annotation with scalable automated data synthesis across demonstration data D={(xi,yi)}\mathcal{D} = \{(x_i, y_i)\} and pairwise preference data P={(xi,yi+,yi−)}\mathcal{P} = \{(x_i, y_i^+, y_i^-)\}:

    1. Ontology Extraction and Instruction Evolution: Tagging is performed using InsTag to construct an instruction ontology. Existing instructions are evolved by prompting language models to inject additional constraints and requirements, producing multi-level task complexity.
    2. Rejection Sampling: For mathematics and reasoning tasks with verifiable ground truths, models sample multiple reasoning paths. Paths reaching verified correct answers serve as demonstration targets (yiy_i), while pairs of correct and incorrect paths form preference pairs (yi+,yi−)(y_i^+, y_i^-).
    3. Execution Feedback and Constraint Verification: For coding tasks, candidate solutions are executed against test cases. For constrained instruction following, models generate Python verification functions that execute against responses to automatically validate formatting and rule compliance.
    4. Data Repurposing: Literary and roleplay instructions are synthesized from open-domain text and structured character profiles (e.g., Wikipedia) using reading comprehension style prompting.
    5. Constitutional Feedback: A constitution of safety principles is used to prompt models to produce explicitly aligned (yi+y_i^+) or violating (yi−y_i^-) responses for safety training.
  5. Knowl 5 — Supervised Fine-Tuning and Two-Stage RLHF Alignment Protocol

    model/method

    The alignment of Qwen2 models follows a sequential post-training protocol:

    1. Supervised Fine-Tuning (SFT): Conducted on over 500,000 instruction-response pairs for 2 epochs with a maximum sequence length of 32,768 tokens. Optimization uses an initial learning rate of 7×10−67 \times 10^{-6} decaying linearly to 7×10−77 \times 10^{-7}, a weight decay of 0.1, and gradient norm clipping at 1.0.
    2. Offline Preference Optimization: Direct Preference Optimization (DPO) is applied over a pre-compiled preference dataset P={(xi,yi+,yi−)}\mathcal{P} = \{(x_i, y_i^+, y_i^-)\} to maximize the relative likelihood of preferred responses yi+y_i^+ over non-preferred responses yi−y_i^-.
    3. Online Reinforcement Learning (Online RLHF): The policy generates multiple responses per prompt in real-time episodes. An online reward model identifies the highest- and lowest-reward responses to construct dynamic preference pairs for iterative DPO updates. The Online Merging Optimizer is applied across checkpoint weights to mitigate the alignment tax (loss of fundamental reasoning capabilities during alignment).
  6. Knowl 6 — Base Model Performance on Standard Benchmarks

    data/table

    Evaluation of the flagship base model Qwen2-72B across English, coding, mathematics, Chinese, and multilingual benchmarks compared with competitive open-weight baselines:

    Datasets Mixtral-8x22B Llama-3-70B Qwen1.5-72B Qwen1.5-110B Qwen2-72B
    English
    MMLU (5-shot) 77.8 79.5 77.5 80.4 84.2
    MMLU-Pro (5-shot) 49.5 52.8 45.8 49.4 55.6
    GPQA (5-shot) 34.3 36.3 36.3 35.9 37.9
    Theorem QA (5-shot) 35.9 32.3 29.3 34.9 43.1
    BBH (3-shot) 78.9 81.0 65.5 74.8 82.4
    HellaSwag (10-shot) 88.7 88.0 86.0 87.5 87.6
    Winogrande (5-shot) 85.0 85.3 83.0 83.5 85.1
    ARC-C (25-shot) 70.7 68.8 65.9 69.6 68.9
    TruthfulQA (0-shot) 51.0 45.6 59.6 49.6 54.8
    Coding
    HumanEval (0-shot) 46.3 48.2 46.3 54.3 64.6
    MBPP (0-shot) 71.7 70.4 66.9 70.9 76.9
    EvalPlus (0-shot) 54.1 54.8 52.9 57.7 65.4
    MultiPL-E (0-shot) 46.7 46.3 41.8 52.7 59.6
    Mathematics
    GSM8K (5-shot) 83.7 83.0 79.5 85.4 89.5
    MATH (4-shot) 41.7 42.5 34.1 49.6 51.1
    Chinese
    C-Eval (5-shot) 54.6 65.2 84.1 89.1 91.0
    CMMLU (5-shot) 53.4 67.2 83.5 88.3 90.1
    Multilingual
    Exam 63.5 70.0 66.4 75.6 76.6
    Understanding 77.7 79.9 78.2 78.2 80.7
    Mathematics (MGSM) 62.9 67.1 61.7 64.4 76.0
    Translation (Flores-101) 23.3 38.0 35.6 36.2 37.8

    Qwen2-72B outperforms Llama-3-70B on MMLU (+4.7), MMLU-Pro (+2.8), GPQA (+1.6), Theorem QA (+9.8), HumanEval (+16.4), and GSM8K (+6.5), while demonstrating significant advantages over previous Qwen1.5 models.

  7. Knowl 7 — Instruction-Tuned 70B+ Model Benchmark Performance

    data/table

    Evaluation of Qwen2-72B-Instruct against state-of-the-art instruction-tuned models across core academic disciplines and alignment/preference benchmarks:

    Datasets Mixtral-8x22B-Inst. Llama-3-70B-Inst. Qwen1.5-72B-Chat Qwen1.5-110B-Chat Qwen2-72B-Inst.
    MMLU 74.0 82.0 75.6 76.5 82.3
    MMLU-Pro 56.1 56.2 51.7 50.5 64.4
    GPQA 49.7 41.9 39.4 32.8 42.4
    Theorem QA 40.8 42.5 28.8 18.8 44.4
    HumanEval 73.8 81.7 71.3 74.4 86.0
    MBPP 75.9 82.3 71.9 76.4 80.2
    MultiPL-E 61.1 63.4 48.1 55.4 69.2
    LiveCodeBench v1 21.8 29.3 17.9 25.3 35.7
    GSM8K 89.1 93.0 82.7 84.5 93.2
    MATH 47.4 50.4 42.5 42.0 69.0
    MT-Bench 8.66 8.95 8.61 8.88 9.12
    MixEval 82.3 84.0 84.1 85.7 86.7
    Arena-Hard 36.4 41.1 36.1 39.8 48.1
    IFEval (strict-prompt) 67.1 77.3 55.8 57.5 77.6
    AlignBench - 7.42 7.28 7.87 8.27

    Qwen2-72B-Instruct exhibits substantial gains in math (MATH score of 69.0 vs 50.4 for Llama-3-70B-Instruct), competitive coding (LiveCodeBench score of 35.7 vs 29.3), and human preference alignment (MT-Bench 9.12, Arena-Hard 48.1).

  8. Knowl 8 — Long-Context Evaluation on NeedleBench and LV-Eval

    data/table

    Long-context retrieval and multi-hop reasoning performance of Qwen2-7B-Instruct and Qwen2-72B-Instruct on NeedleBench and LV-Eval across context lengths up to 256k tokens, with and without YaRN and Dual Chunk Attention (DCA):

    Model NeedleBench LV-Eval
    8k 32k 128k 256k 16k 32k 64k 128k 256k
    ChatGLM4-9B-1M 56.61 49.15 44.30 45.29 46.40 43.23 42.92 40.41 36.95
    Qwen2-7B-Instruct 87.07 73.64 38.77 2.92 49.77 46.93 28.03 11.01 0.55
    + YARN + DCA 87.07 73.64 66.32 60.71 49.77 46.93 42.14 36.64 34.72
    Qwen2-72B-Instruct 91.90 92.01 73.05 17.13 58.82 56.70 42.92 31.79 2.88
    + YARN + DCA 91.90 92.01 90.27 85.21 58.82 56.70 53.03 48.83 42.35

    Within 32k tokens, the models operate natively. Beyond 32k tokens, adding YaRN and DCA prevents catastrophic performance degradation: Qwen2-72B-Instruct maintains a NeedleBench score of 85.21 and LV-Eval score of 42.35 at 256k tokens (compared to 17.13 and 2.88 without YaRN+DCA).

  9. Knowl 9 — Multilingual and Safety Evaluation of Qwen2-72B-Instruct

    data/table

    Human evaluation of multilingual generation across 10 languages (rated 1 to 5 by professional native-language annotators) and harmful prompt rejection rate on safety benchmarks (lower percentage indicates fewer harmful responses generated):

    Language (Score 1–5) GPT-3.5-Turbo GPT-4-Turbo GPT-4o Claude-3-Opus Qwen2-72B-Inst.
    Arabic 2.52 3.44 3.55 4.15 3.86
    French 3.47 4.19 4.16 4.23 4.01
    Indonesian 3.56 4.09 4.39 4.40 3.83
    Japanese 2.75 3.68 3.72 3.85 3.63
    Korean 2.37 4.24 4.40 4.23 4.14
    Portuguese 3.37 3.86 3.89 4.09 3.97
    Russian 3.24 4.27 4.32 4.25 4.15
    Spanish 4.07 4.08 4.26 4.31 4.10
    Thai 3.38 4.11 4.09 4.01 3.75
    Vietnamese 3.90 3.84 4.14 3.98 3.91
    Average 3.16 3.98 4.09 4.15 3.93
    Risk Category (% Harmful) GPT-4 Mixtral-8x22B Qwen2-72B-Inst.
    Illegal 0.00 6.87 0.00
    Fraud 3.40 8.49 2.41
    Pornography 23.63 33.82 22.91
    Privacy 3.37 15.03 2.47

    Qwen2-72B-Instruct achieves an average multilingual score of 3.93, approaching GPT-4-Turbo (3.98), while achieving lower harmful response rates on fraud (2.41%), privacy (2.47%), and pornography (22.91%) compared to GPT-4.

  10. Knowl 10 — Training Data Decontamination Protocol via LCS and N-Gram Filtering

    model/method

    To prevent benchmark contamination without excessively discarding false positives caused by ubiquitous coding syntax and mathematical formulas, Qwen2 uses a dual-criterion filtering protocol on the pre-training and post-training data against evaluation sets:

    1. Text Normalization: All symbols, punctuation, and formatting are stripped from both training sequences sts_t and evaluation sequences ses_e, followed by tokenization.
    2. Longest Common Subsequence (LCS) Constraint: A training sequence sts_t is identified as contaminated and filtered out if there exists an evaluation sequence ses_e satisfying:

    ∣LCS(st,se)∣≥13and∣LCS(st,se)∣≥0.6×min⁡(∣st∣,∣se∣)|\text{LCS}(s_t, s_e)| \ge 13 \quad \text{and} \quad |\text{LCS}(s_t, s_e)| \ge 0.6 \times \min(|s_t|, |s_e|)

    where ∣LCS(st,se)∣|\text{LCS}(s_t, s_e)| is the length of the longest common subsequence of tokens between sts_t and ses_e.

    In validation using a strict 13-gram overlap test without the LCS ratio constraint, removing contaminated samples yielded minimal change on downstream benchmarks (e.g., MMLU changed by +0.9 on Qwen2-72B-Instruct and +0.8 on Qwen2-7B-Instruct), indicating that detected overlaps were dominated by generic formulas and boilerplate rather than verbatim task leakage.

Coverage note — Omitted individual baseline result tables for smaller models (Tables 3, 4, 5, 7, 8, 9) and internal automatic evaluations (Tables 10, 11) to focus on the definitive representative benchmark comparisons (Tables 2, 6, 12, 13, 14) and all architectural, training, and alignment methodologies.

References

  1. 1.Marah Abdin, Jyoti Aneja, Sebastien Bubeck, Caio Cesar Teodoro Mendes, Weizhu Chen, Allie Del Giorno, Ronen Eldan, Sivakanth Gopi, Suriya Gunasekar, Mojan Javaheripi, Piero Kauffmann, Yin Tat Lee, Yuanzhi Li, Anh Nguyen, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Michael Santacroce, Harkirat Singh Behl, Adam Taumann Kalai, Xin Wang, Rachel Ward, Philipp Witte, Cyril Zhang, and Yi Zhang. Phi-2: The surprising power of small language models, 2024. URL https://www.microsoft.com/en-us/research/blog/phi-2-the-surprising-power-of-small-language-models/.
  2. 2.AI@Meta. Llama 3 model card, 2024. URL https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md.
  3. 3.Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. GQA: Training generalized multi-query Transformer models from multi-head checkpoints. In EMNLP, pp. 4895–4901. Association for Computational Linguistics, 2023.
  4. 4.Chenxin An, Fei Huang, Jun Zhang, Shansan Gong, Xipeng Qiu, Chang Zhou, and Lingpeng Kong. Training-free long-context scaling of large language models. CoRR, abs/2402.17463, 2024.
  5. 5.Anthropic. The Claude 3 model family: Opus, Sonnet, Haiku. Technical report, Anthropic, AI, 2024. URL https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf.
  6. 6.Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. Program synthesis with large language models. CoRR, abs/2108.07732, 2021.
  7. 7.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. Qwen technical report. CoRR, abs/2309.16609, 2023a.
  8. 8.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A frontier large vision-language model with versatile abilities. CoRR, abs/2308.12966, 2023b.
  9. 9.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosiute, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemí Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. Constitutional AI: Harmlessness from AI feedback. CoRR, abs/2212.08073, 2022.
  10. 10.Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. The Belebele benchmark: A parallel reading comprehension dataset in 122 language variants. CoRR, abs/2308.16884, 2023.
  11. 11.Boxi Cao, Keming Lu, Xinyu Lu, Jiawei Chen, Mengjie Ren, Hao Xiang, Peilin Liu, Yaojie Lu, Ben He, Xianpei Han, Le Sun, Hongyu Lin, and Bowen Yu. Towards scalable automated alignment of LLMs: A survey. CoRR, abs/2406.01252, 2024.
  12. 12.Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q. Feldman, Arjun Guha, Michael Greenberg, and Abhinav Jangda. MultiPL-E: A scalable and polyglot approach to benchmarking neural code generation. IEEE Trans. Software Eng., 49(7):3675–3691, 2023.
  13. 13.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code. CoRR, abs/2107.03374, 2021.
  14. 14.Wenhu Chen, Ming Yin, Max Ku, Pan Lu, Yixin Wan, Xueguang Ma, Jianyu Xu, Xinyi Wang, and Tony Xia. TheoremQA: A theorem-driven question answering dataset. In EMNLP, pp. 7889–7901. Association for Computational Linguistics, 2023a.
  15. 15.Zhihong Chen, Shuo Yan, Juhao Liang, Feng Jiang, Xiangbo Wu, Fei Yu, Guiming Hardy Chen, Junying Chen, Hongbo Zhang, Li Jianquan, Wan Xiang, and Benyou Wang. Multilingual-SIFT: Multilingual supervised instruction fine-tuning, 2023b. URL https://github.com/FreedomIntelligence/MultilingualSIFT.
  16. 16.Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael I. Jordan, Joseph E. Gonzalez, and Ion Stoica. Chatbot arena: An open platform for evaluating LLMs by human preference. CoRR, abs/2403.04132, 2024.
  17. 17.Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. Qwen-Audio: Advancing universal audio understanding via unified large-scale audio-language models. CoRR, abs/2311.07919, 2023.
  18. 18.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? Try ARC, the AI2 reasoning challenge. CoRR, abs/1803.05457, 2018.
  19. 19.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. CoRR, abs/2110.14168, 2021.
  20. 20.Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models. CoRR, abs/2401.06066, 2024.
  21. 21.Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier. Language modeling with gated convolutional networks. In ICML, volume 70 of Proceedings of Machine Learning Research, pp. 933–941. PMLR, 2017.
  22. 22.Guanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li, Mingfeng Xue, Dayiheng Liu, Wei Wang, Zheng Yuan, Chang Zhou, and Jingren Zhou. How abilities in large language models are affected by supervised fine-tuning data composition. CoRR, abs/2310.05492, 2023.
  23. 23.Guanting Dong, Keming Lu, Chengpeng Li, Tingyu Xia, Bowen Yu, Chang Zhou, and Jingren Zhou. Self-play with execution feedback: Improving instruction-following capabilities of large language models. CoRR, abs/2406.13542, 2024.
  24. 24.Alena Fenogenova, Artem Chervyakov, Nikita Martynov, Anastasia Kozlova, Maria Tikhonova, Albina Akhmetgareeva, Anton A. Emelyanov, Denis Shevelev, Pavel Lebedev, Leonid Sinev, Ulyana Isaeva, Katerina Kolomeytseva, Daniil Moskovskiy, Elizaveta Goncharova, Nikita Savushkin, Polina Mikhailova, Denis Dimitrov, Alexander Panchenko, and Sergey Markov. MERA: A comprehensive LLM evaluation in russian. CoRR, abs/2401.04531, 2024.
  25. 25.Shahriar Golchin and Mihai Surdeanu. Time travel in llms: Tracing data contamination in large language models. In ICLR. OpenReview.net, 2024.
  26. 26.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. The Flores-101 evaluation benchmark for low-resource and multilingual machine translation. Trans. Assoc. Comput. Linguistics, 10:522–538, 2022.
  27. 27.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In ICLR. OpenReview.net, 2021a.
  28. 28.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. In NeurIPS Datasets and Benchmarks, 2021b.
  29. 29.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. C-Eval: A multi-level multi-discipline chinese evaluation suite for foundation models. In NeurIPS, 2023.
  30. 30.Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. LiveCodeBench: Holistic and contamination free evaluation of large language models for code. CoRR, abs/2403.07974, 2024.
  31. 31.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7B. CoRR, abs/2310.06825, 2023a.
  32. 32.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mixtral of experts. CoRR, abs/2401.04088, 2024.
  33. 33.Zixuan Jiang, Jiaqi Gu, Hanqing Zhu, and David Z. Pan. Pre-RMSNorm and Pre-CRMSNorm Transformers: Equivalent and efficient pre-LN Transformers. CoRR, abs/2305.14858, 2023b.
  34. 34.Gregory Kamradt. Needle in a haystack - pressure testing LLMs, 2023. URL https://github.com/gkamradt/LLMTest_NeedleInAHaystack.
  35. 35.Aran Komatsuzaki, Joan Puigcerver, James Lee-Thorp, Carlos Riquelme Ruiz, Basil Mustafa, Joshua Ainslie, Yi Tay, Mostafa Dehghani, and Neil Houlsby. Sparse upcycling: Training mixture-of-experts from dense checkpoints. In ICLR. OpenReview.net, 2023.
  36. 36.Fajri Koto, Nurul Aisyah, Haonan Li, and Timothy Baldwin. Large language models only pass primary school exams in Indonesia: A comprehensive test on IndoMMLU. In EMNLP, pp. 12359–12374. Association for Computational Linguistics, 2023.
  37. 37.Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. CMMLU: Measuring massive multitask language understanding in Chinese. CoRR, abs/2306.09212, 2023.
  38. 38.Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. From crowdsourced data to high-quality benchmarks: Arena-Hard and BenchBuilder pipeline. CoRR, abs/2406.11939, 2024.
  39. 39.Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, Omri Abend, Raz Alon, Tomer Asida, Amir Bergman, Roman Glozman, Michael Gokhman, Avashalom Manevich, Nir Ratner, Noam Rozen, Erez Shwartz, Mor Zusman, and Yoav Shoham. Jamba: A hybrid Transformer-Mamba language model. CoRR, abs/2403.19887, 2024.
  40. 40.Stephanie Lin, Jacob Hilton, and Owain Evans. TruthfulQA: Measuring how models mimic human falsehoods. In ACL (1), pp. 3214–3252. Association for Computational Linguistics, 2022a.
  41. 41.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona T. Diab, Veselin Stoyanov, and Xian Li. Few-shot learning with multilingual generative language models. In EMNLP, pp. 9019–9052. Association for Computational Linguistics, 2022b.
  42. 42.Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. Is your code generated by ChatGPT really correct? Rigorous evaluation of large language models for code generation. In NeurIPS, 2023a.
  43. 43.Xiao Liu, Xuanyu Lei, Shengyuan Wang, Yue Huang, Zhuoer Feng, Bosi Wen, Jiale Cheng, Pei Ke, Yifan Xu, Weng Lam Tam, Xiaohan Zhang, Lichao Sun, Hongning Wang, Jing Zhang, Minlie Huang, Yuxiao Dong, and Jie Tang. AlignBench: Benchmarking Chinese alignment of large language models. CoRR, abs/2311.18743, 2023b.
  44. 44.Keming Lu, Bowen Yu, Fei Huang, Yang Fan, Runji Lin, and Chang Zhou. Online merging optimizers for boosting rewards and mitigating tax in alignment. CoRR, abs/2405.17931, 2024a.
  45. 45.Keming Lu, Bowen Yu, Chang Zhou, and Jingren Zhou. Large language models are superpositions of all characters: Attaining arbitrary role-play via self-alignment. CoRR, abs/2401.12474, 2024b.
  46. 46.Keming Lu, Hongyi Yuan, Zheng Yuan, Runji Lin, Junyang Lin, Chuanqi Tan, Chang Zhou, and Jingren Zhou. #InsTag: Instruction tagging for analyzing supervised fine-tuning of large language models. In ICLR. OpenReview.net, 2024c.
  47. 47.Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Riviere, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Clément Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, Justin Mao-Jones, Katherine Lee, Kathy Yu, Katie Millican, Lars Lowe Sjoesund, Lisa Lee, Lucas Dixon, Machel Reid, Maciej Mikuła, Mateo Wirth, Michael Sharman, Nikolai Chinaev, Nithum Thain, Olivier Bachem, Oscar Chang, Oscar Wahltinez, Paige Bailey, Paul Michel, Petko Yotov, Rahma Chaabouni, Ramona Comanescu, Reena Jana, Rohan Anil, Ross McIlroy, Ruibo Liu, Ryan Mullins, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sholto Douglas, Shree Pandya, Siamak Shakeri, Soham De, Ted Klimenko, Tom Hennigan, Vlad Feinberg, Wojciech Stokowiec, Yu hui Chen, Zafarali Ahmed, Zhitao Gong, Tris Warkentin, Ludovic Peran, Minh Giang, Clément Farabet, Oriol Vinyals, Jeff Dean, Koray Kavukcuoglu, Demis Hassabis, Zoubin Ghahramani, Douglas Eck, Joelle Barral, Fernando Pereira, Eli Collins, Armand Joulin, Noah Fiedel, Evan Senter, Alek Andreev, and Kathleen Kenealy. Gemma: Open models based on Gemini research and technology. CoRR, abs/2403.08295, 2024.
  48. 48.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M. Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. Crosslingual generalization through multitask finetuning. In ACL (1), pp. 15991–16111. Association for Computational Linguistics, 2023.
  49. 49.Jinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng, Mahir Shah, Kabir Jain, Graham Neubig, and Yang You. MixEval: Deriving wisdom of the crowd from LLM benchmark mixtures. CoRR, abs/2406.06565, 2024.
  50. 50.OpenAI. Introducing ChatGPT, 2022. URL https://openai.com/index/chatgpt/.
  51. 51.OpenAI. GPT4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  52. 52.OpenAI. Hello GPT-4o, 2024. URL https://openai.com/index/hello-gpt-4o/.
  53. 53.OpenCompass Contributors. OpenCompass: A universal evaluation platform for foundation models, 2023. URL https://github.com/open-compass/opencompass.
  54. 54.Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. YaRN: Efficient context window extension of large language models. CoRR, abs/2309.00071, 2023.
  55. 55.Edoardo Maria Ponti, Goran Glavas, Olga Majewska, Qianchu Liu, Ivan Vulic, and Anna Korhonen. XCOPA: A multilingual dataset for causal commonsense reasoning. In EMNLP (1), pp. 2362–2376. Association for Computational Linguistics, 2020.
  56. 56.Qwen Team. Introducing Qwen1.5, 2024a. URL https://qwenlm.github.io/blog/qwen1.5/.
  57. 57.Qwen Team. Qwen1.5-110B: The first 100B+ model of the Qwen1.5 series, 2024b. URL https://qwenlm.github.io/blog/qwen1.5-110b/.
  58. 58.Qwen Team. Qwen1.5-MoE: Matching 7B model performance with 1/3 activated parameters, 2024c. URL https://qwenlm.github.io/blog/qwen-moe/.
  59. 59.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In NeurIPS, 2023.
  60. 60.Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He. DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation AI scale. In ICML, volume 162 of Proceedings of Machine Learning Research, pp. 18332–18346. PMLR, 2022.
  61. 61.Mathieu Ravaut, Bosheng Ding, Fangkai Jiao, Hailin Chen, Xingxuan Li, Ruochen Zhao, Chengwei Qin, Caiming Xiong, and Shafiq Joty. How much are LLMs contaminated? A comprehensive survey and the llmsanitize library. CoRR, abs/2404.00699, 2024.
  62. 62.David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level Google-proof Q&A benchmark. CoRR, abs/2311.12022, 2023.
  63. 63.Oscar Sainz, Jon Ander Campos, Iker García-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre. NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark. In EMNLP (Findings), pp. 10776–10787. Association for Computational Linguistics, 2023.
  64. 64.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. WinoGrande: An adversarial winograd schema challenge at scale. Commun. ACM, 64(9):99–106, 2021.
  65. 65.Jianlin Su. The magical effect of the Bias term: RoPE + Bias = better length extrapolation, 2023. URL https://spaces.ac.cn/archives/9577.
  66. 66.Jianlin Su, Murtadha H. M. Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced Transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
  67. 67.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. Challenging BIG-Bench tasks and whether chain-of-thought can solve them. In ACL (Findings), pp. 13003–13051. Association for Computational Linguistics, 2023.
  68. 68.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. LLaMA: Open and efficient foundation language models. CoRR, abs/2302.13971, 2023.
  69. 69.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, pp. 5998–6008, 2017.
  70. 70.Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. MMLU-Pro: A more robust and challenging multi-task language understanding benchmark. CoRR, abs/2406.01574, 2024.
  71. 71.Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, and Hao Ma. Effective long-context scaling of foundation models. CoRR, abs/2309.16039, 2023.
  72. 72.Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. PAWS-X: A cross-lingual adversarial dataset for paraphrase identification. In EMNLP/IJCNLP (1), pp. 3685–3690. Association for Computational Linguistics, 2019.
  73. 73.Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and Zonghong Dai. Yi: Open foundation models by 01.AI. CoRR, abs/2403.04652, 2024.
  74. 74.Tao Yuan, Xuefei Ning, Dong Zhou, Zhijie Yang, Shiyao Li, Minghui Zhuang, Zheyue Tan, Zhuyu Yao, Dahua Lin, Boxun Li, Guohao Dai, Shengen Yan, and Yu Wang. LV-Eval: A balanced long-context benchmark with 5 length levels up to 256K. CoRR, abs/2402.05136, 2024.
  75. 75.Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Chuanqi Tan, and Chang Zhou. Scaling relationship on learning mathematical reasoning with large language models. CoRR, abs/2308.01825, 2023.
  76. 76.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In ACL (1), pp. 4791–4800. Association for Computational Linguistics, 2019.
  77. 77.Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong, Mingdao Liu, Minlie Huang, Peng Zhang, Qinkai Zheng, Rui Lu, Shuaiqi Duan, Shudan Zhang, Shulin Cao, Shuxun Yang, Weng Lam Tam, Wenyi Zhao, Xiao Liu, Xiao Xia, Xiaohan Zhang, Xiaotao Gu, Xin Lv, Xinghan Liu, Xinyi Liu, Xinyue Yang, Xixuan Song, Xunkai Zhang, Yifan An, Yifan Xu, Yilin Niu, Yuantao Yang, Yueyan Li, Yushi Bai, Yuxiao Dong, Zehan Qi, Zhaoyu Wang, Zhen Yang, Zhengxiao Du, Zhenyu Hou, and Zihan Wang. ChatGLM: A family of large language models from GLM-130B to GLM-4 all tools. CoRR, abs/2406.12793, 2024.
  78. 78.Yingxiu Zhao, Bowen Yu, Binyuan Hui, Haiyang Yu, Minghao Li, Fei Huang, Nevin L. Zhang, and Yongbin Li. Tree-Instruct: A preliminary study of the intrinsic relationship between complexity and alignment. In LREC/COLING, pp. 16776–16789. ELRA and ICCL, 2024.
  79. 79.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In NeurIPS, 2023.
  80. 80.Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-following evaluation for large language models. CoRR, abs/2311.07911, 2023.

Citation

MLA
Yang, A., et al. “Qwen2 Technical Report”. arXiv, 2024, http://arxiv.org/abs/2407.10671v4.
APA
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Tang, J., Wang, J., Yang, J., Tu, J., Zhang, J., Ma, J., … Fan, Z. (2024). Qwen2 Technical Report. arXiv. http://arxiv.org/abs/2407.10671v4
Chicago
Yang, A., B. Yang, B. Hui, et al. 2024. “Qwen2 Technical Report”. arXiv. http://arxiv.org/abs/2407.10671v4.
Harvard
Yang, A. et al. (2024) “Qwen2 Technical Report”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2407.10671v4.
Vancouver
1. Yang A, Yang B, Hui B, et al (2024) Qwen2 Technical Report. arXiv

BibTeX

@article{yang2024qwen2,
  title = {Qwen2 Technical Report},
  author = {Yang, An and Yang, Baosong and Hui, Binyuan and Zheng, Bo and Yu, Bowen and Zhou, Chang and Li, Chengpeng and Li, Chengyuan and Liu, Dayiheng and Huang, Fei and Dong, Guanting and Wei, Haoran and Lin, Huan and Tang, Jialong and Wang, Jialin and Yang, Jian and Tu, Jianhong and Zhang, Jianwei and Ma, Jianxin and Yang, Jianxin and Xu, Jin and Zhou, Jingren and Bai, Jinze and He, Jinzheng and Lin, Junyang and Dang, Kai and Lu, Keming and Chen, Keqin and Yang, Kexin and Li, Mei and Xue, Mingfeng and Ni, Na and Zhang, Pei and Wang, Peng and Peng, Ru and Men, Rui and Gao, Ruize and Lin, Runji and Wang, Shijie and Bai, Shuai and Tan, Sinan and Zhu, Tianhang and Li, Tianhao and Liu, Tianyu and Ge, Wenbin and Deng, Xiaodong and Zhou, Xiaohuan and Ren, Xingzhang and Zhang, Xinyu and Wei, Xipin and Ren, Xuancheng and Liu, Xuejing and Fan, Yang and Yao, Yang and Zhang, Yichang and Wan, Yu and Chu, Yunfei and Liu, Yuqiong and Cui, Zeyu and Zhang, Zhenru and Guo, Zhifang and Fan, Zhihao},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2407.10671v4},
  eprint = {2407.10671}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors