DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Zhihong ShaoPeiyi WangQihao ZhuRunxin XuJun-Mei SongMingchuan ZhangY. K. LiYu WuDaya Guo

article2024arXiv8,752 citations

Introduces DeepSeekMath 7B and Group Relative Policy Optimization (GRPO), demonstrating that an open 7B-parameter model can approach closed frontier performance on the competition-level MATH benchmark through large-scale pre-training and memory-efficient reinforcement learning.

Listen

Researchers at DeepSeek-AI developed DeepSeekMath 7B to close the performance gap between open and closed language models on mathematical reasoning tasks. The work addresses the fact that leading systems such as GPT-4 and Gemini-Ultra remain unavailable while existing open models lag well behind on competition-level benchmarks. The authors created a 120-billion-token mathematics corpus from Common Crawl through an iterative fastText classifier trained on high-quality seeds and refined with human annotation. They continued pre-training DeepSeek-Coder-Base-v1.5 7B on this corpus plus code and natural-language data, applied targeted instruction tuning, and then introduced Group Relative Policy Optimization, a memory-efficient reinforcement-learning variant that eliminates the separate value model used in standard Proximal Policy Optimization.

The resulting DeepSeekMath-Base 7B already matched or exceeded Minerva 540B on GSM8K and MATH while using roughly one-eightieth the parameters. After instruction tuning and GRPO, the final model reached 51.7 percent on MATH without tools or voting and 60.9 percent with self-consistency over 64 samples, surpassing all prior open 7B-to-70B models and most proprietary systems. Code pre-training before mathematics data improved both tool-using and tool-free reasoning, whereas arXiv papers produced negligible gains. GRPO delivered consistent lifts on both in-domain and out-of-domain tasks while cutting memory requirements, and the gains appear to come from sharpening the output distribution rather than expanding core capabilities.

These results show that publicly available web data, when filtered rigorously, can support frontier-level mathematical performance at modest scale, and that simplified reinforcement learning can further improve already strong instruction-tuned models without additional labeled data. Organizations seeking reliable quantitative reasoning therefore have a practical path to deploy capable open models rather than relying solely on closed APIs.

Further progress will require stronger process-level reward models, sampling strategies that move beyond nucleus sampling, and explicit handling of noisy reward signals. The authors also note that geometry and formal-proof performance remain weaker than closed models and that few-shot gains are limited by current scale. Continued refinement of the data pipeline and exploration of robust weak-to-strong alignment methods are the most direct next steps supported by the evidence.

  • Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Reading this foundational introduction to chain-of-thought prompting is essential because the source paper builds directly upon step-by-step reasoning mechanisms to enhance mathematical problem-solving.
  • Paper: Measuring Mathematical Problem Solving With the MATH Dataset, Dan Hendrycks et al. (2021). Familiarity with this competition-level math benchmark is required to understand the primary evaluation domain addressed by the source paper.
  • Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). This paper's exploration of reward modeling and verification steps provides vital context for understanding the reinforcement learning techniques evaluated in the source.
Cover for DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Abstract

Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pre-training DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens sourced from Common Crawl, together with natural language and code data. DeepSeekMath 7B has achieved an impressive score of 51.7% on the competition-level MATH benchmark without relying on external toolkits and voting techniques, approaching the performance level of Gemini-Ultra and GPT-4. Self-consistency over 64 samples from DeepSeekMath 7B achieves 60.9% on MATH. The mathematical reasoning capability of DeepSeekMath is attributed to two key factors: First, we harness the significant potential of publicly available web data through a meticulously engineered data selection pipeline. Second, we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.

Table of Contents

  • 1 Introduction
  • 1.1 Contributions
  • 1.2 Summary of Evaluations and Metrics
  • 2 Math Pre-Training
  • 2.1 Data Collection and Decontamination
  • 2.2 Validating the Quality of the DeepSeekMath Corpus
  • 2.2.1 Training Setting
  • 2.2.2 Evaluation Results
  • 2.3 Training and Evaluating DeepSeekMath-Base 7B
  • 3 Supervised Fine-Tuning
  • 3.1 SFT Data Curation
  • 3.2 Training and Evaluating DeepSeekMath-Instruct 7B
  • 4 Reinforcement Learning
  • 4.1 Group Relative Policy Optimization
  • 4.1.1 From PPO to GRPO
  • 4.1.2 Outcome Supervision RL with GRPO
  • 4.1.3 Process Supervision RL with GRPO
  • 4.1.4 Iterative RL with GRPO
  • 4.2 Training and Evaluating DeepSeekMath-RL
  • 5 Discussion
  • 5.1 Lessons Learnt in Pre-Training
  • 5.1.1 Code Training Benefits Mathematical Reasoning
  • 5.1.2 ArXiv Papers Seem Ineffective in Improving Mathematical Reasoning
  • 5.2 Insights of Reinforcement Learning
  • 5.2.1 Towards to a Unified Paradigm
  • 5.2.2 Why RL Works?
  • 5.2.3 How to Achieve More Effective RL?
  • 6 Conclusion, Limitation, and Future Work
  • References
  • A Appendix
  • A.1 Analysis of Reinforcement Learning
  • A.1.1 Supervised Fine-tuning
  • A.1.2 Rejection Sampling Fine-tuning
  • A.1.3 Online Rejection Sampling Fine-tuning
  • A.1.4 Direct Preference Optimization (DPO)
  • A.1.5 Proximal Policy Optimization (PPO)
  • A.1.6 Group Relative Policy Optimization (GRPO)

Knowls

  1. Knowl 1 — Group Relative Policy Optimization

    model/method

    Group Relative Policy Optimization (GRPO) is a reinforcement learning algorithm designed to align large language models without instantiating an auxiliary critic (value) model, thereby substantially reducing memory and compute overhead compared to Proximal Policy Optimization (PPO).

    For each question qq sampled from a prompt distribution P(Q)P(Q), GRPO samples a group of GG distinct outputs {o1,o2,,oG}\{o_1, o_2, \dots, o_G\} from the previous policy πθold\pi_{\theta_{\mathrm{old}}}. The policy model πθ\pi_\theta is optimized by maximizing the following surrogate objective:

    JGRPO(θ)=E[qP(Q),{oi}i=1Gπθold(Oq)][1Gi=1G1oit=1oi(min(πθ(oi,tq,oi,<t)πθold(oi,tq,oi,<t)A^i,t,clip(πθ(oi,tq,oi,<t)πθold(oi,tq,oi,<t),1ε,1+ε)A^i,t)βDKL(πθπref))]J_{\mathrm{GRPO}}(\theta) = \mathbb{E}\left[q \sim P(Q), \{o_i\}_{i=1}^G \sim \pi_{\theta_{\mathrm{old}}}(O|q)\right] \left[ \frac{1}{G} \sum_{i=1}^G \frac{1}{|o_i|} \sum_{t=1}^{|o_i|} \left( \min\left( \frac{\pi_\theta(o_{i,t} | q, o_{i,<t})}{\pi_{\theta_{\mathrm{old}}}(o_{i,t} | q, o_{i,<t})} \hat{A}_{i,t}, \operatorname{clip}\left( \frac{\pi_\theta(o_{i,t} | q, o_{i,<t})}{\pi_{\theta_{\mathrm{old}}}(o_{i,t} | q, o_{i,<t})}, 1-\varepsilon, 1+\varepsilon \right) \hat{A}_{i,t} \right) - \beta D_{\mathrm{KL}}(\pi_\theta \parallel \pi_{\mathrm{ref}}) \right) \right]

    where ε\varepsilon is a clipping threshold, β\beta is the Kullback-Leibler (KL) divergence penalty coefficient, πref\pi_{\mathrm{ref}} is a frozen reference policy (typically the supervised fine-tuned model), and A^i,t\hat{A}_{i,t} is the token-level advantage computed from the relative normalized rewards within the sampled group {o1,,oG}\{o_1, \dots, o_G\}.

    The per-token KL divergence between the current policy πθ\pi_\theta and the reference policy πref\pi_{\mathrm{ref}} is computed using the unbiased, non-negative estimator:

    DKL(πθπref)=πref(oi,tq,oi,<t)πθ(oi,tq,oi,<t)logπref(oi,tq,oi,<t)πθ(oi,tq,oi,<t)1D_{\mathrm{KL}}(\pi_\theta \parallel \pi_{\mathrm{ref}}) = \frac{\pi_{\mathrm{ref}}(o_{i,t} | q, o_{i,<t})}{\pi_\theta(o_{i,t} | q, o_{i,<t})} - \log \frac{\pi_{\mathrm{ref}}(o_{i,t} | q, o_{i,<t})}{\pi_\theta(o_{i,t} | q, o_{i,<t})} - 1

  2. Knowl 2 — Iterative Group Relative Policy Optimization Algorithm

    algorithm

    Iterative Group Relative Policy Optimization alternately samples reasoning trajectories, computes group-relative advantages, updates the actor policy via the clipped objective, and updates the reward model on fresh on-policy generations combined with historical replay data.

    Input: initial policy model πθinit\pi_{\theta_{\mathrm{init}}}, reward model rφr_\varphi, prompt dataset D\mathcal{D}, hyperparameters ε,β,μ\varepsilon, \beta, \mu, total iterations II, steps per iteration MM, group size GG
    πθπθinit\pi_\theta \leftarrow \pi_{\theta_{\mathrm{init}}}
    for iteration = 1 to II do
        πrefπθ\pi_{\mathrm{ref}} \leftarrow \pi_\theta
        for step = 1 to MM do
            Sample a prompt batch DbD\mathcal{D}_b \sim \mathcal{D}
            πθoldπθ\pi_{\theta_{\mathrm{old}}} \leftarrow \pi_\theta
            for each prompt qDbq \in \mathcal{D}_b do
                Sample GG completions {oi}i=1Gπθold(q)\{o_i\}_{i=1}^G \sim \pi_{\theta_{\mathrm{old}}}(\cdot | q)
                Compute scalar/step rewards {ri}i=1G\{r_i\}_{i=1}^G by scoring {oi}i=1G\{o_i\}_{i=1}^G with rφr_\varphi
                Compute per-token advantages A^i,t\hat{A}_{i,t} using group-relative reward normalization
            for GRPO step = 1 to μ\mu do
                Update policy πθ\pi_\theta by maximizing the GRPO surrogate objective
            Update reward model rφr_\varphi via continuous training on newly sampled policy data mixed with 10% historical replay data
    Output: trained policy πθ\pi_\theta
  3. Knowl 3 — Outcome and Process Advantage Estimation in GRPO

    model/method

    In Group Relative Policy Optimization (GRPO), advantage estimation eliminates learned value baselines by normalizing rewards within a group of GG sampled outputs {o1,,oG}\{o_1, \dots, o_G\} for a prompt qq under either outcome or process supervision.

    Under Outcome Supervision (OS), an outcome reward model assigns a single scalar reward rir_i to each complete sequence oio_i. The group reward vector r={r1,r2,,rG}\mathbf{r} = \{r_1, r_2, \dots, r_G\} is normalized as:

    r~i=rimean(r)std(r)\tilde{r}_i = \frac{r_i - \operatorname{mean}(\mathbf{r})}{\operatorname{std}(\mathbf{r})}

    The advantage assigned to every token t{1,,oi}t \in \{1, \dots, |o_i|\} across the output sequence is uniform:

    A^i,t=r~i\hat{A}_{i,t} = \tilde{r}_i

    Under Process Supervision (PS), a process reward model evaluates each reasoning step j{1,,Ki}j \in \{1, \dots, K_i\} in sequence oio_i, producing step rewards riindex(j)r_i^{\mathrm{index}(j)} where index(j)\mathrm{index}(j) represents the token index at the end of step jj. The collection of all step rewards in the group R={{riindex(j)}j=1Ki}i=1GR = \{\{r_i^{\mathrm{index}(j)}\}_{j=1}^{K_i}\}_{i=1}^G is normalized collectively:

    r~iindex(j)=riindex(j)mean(R)std(R)\tilde{r}_i^{\mathrm{index}(j)} = \frac{r_i^{\mathrm{index}(j)} - \operatorname{mean}(R)}{\operatorname{std}(R)}

    The advantage for token tt in output oio_i is computed by accumulating all normalized rewards from subsequent steps:

    A^i,t=index(j)tr~iindex(j)\hat{A}_{i,t} = \sum_{\mathrm{index}(j) \ge t} \tilde{r}_i^{\mathrm{index}(j)}

  4. Knowl 4 — Unified Policy Gradient Formulation for LLM Alignment

    model/method

    Post-training alignment approaches (Supervised Fine-Tuning, Rejection Sampling Fine-Tuning, Direct Preference Optimization, Proximal Policy Optimization, and Group Relative Policy Optimization) can be unified into a single policy gradient parameter update equation:

    θJA(θ)=E(q,o)D[1ot=1oGCA(q,o,t,πrf)θlogπθ(otq,o<t)]\nabla_\theta J_{\mathcal{A}}(\theta) = \mathbb{E}_{(q, o) \sim \mathcal{D}} \left[ \frac{1}{|o|} \sum_{t=1}^{|o|} GC_{\mathcal{A}}(q, o, t, \pi_{rf}) \nabla_\theta \log \pi_\theta(o_t | q, o_{<t}) \right]

    where D\mathcal{D} denotes the sampling data source, πrf\pi_{rf} is the reward scoring mechanism (rule-based evaluation or learned reward model), and GCAGC_{\mathcal{A}} is the method-specific gradient coefficient:

    • Supervised Fine-Tuning (SFT): (q,o)Psft(Q,O)(q, o) \sim P_{\mathrm{sft}}(Q, O), GCSFT=1GC_{\mathrm{SFT}} = 1.
    • Rejection Sampling Fine-Tuning (RFT): qPsft(Q),oπsft(Oq)q \sim P_{\mathrm{sft}}(Q), o \sim \pi_{\mathrm{sft}}(O|q), GCRFT=I(o is correct)GC_{\mathrm{RFT}} = \mathbb{I}(o \text{ is correct}). Incorrect responses receive GC=0GC = 0, while correct responses receive a uniform +1+1.
    • Online Rejection Sampling Fine-Tuning (Online RFT): qPsft(Q),oπθ(Oq)q \sim P_{\mathrm{sft}}(Q), o \sim \pi_\theta(O|q), GCOnlineRFT=I(o is correct)GC_{\mathrm{Online RFT}} = \mathbb{I}(o \text{ is correct}). Samples are drawn on-policy from the real-time model πθ\pi_\theta.
    • Direct Preference Optimization (DPO): qPsft(Q)q \sim P_{\mathrm{sft}}(Q), paired outputs o+,oπsft(Oq)o^+, o^- \sim \pi_{\mathrm{sft}}(O|q), with gradient coefficient:

    GCDPO(q,o,t)=σ(βlogπθ(otq,o<t)πref(otq,o<t)βlogπθ(ot+q,o<t+)πref(ot+q,o<t+))GC_{\mathrm{DPO}}(q, o, t) = \sigma\left( \beta \log \frac{\pi_\theta(o^-_t | q, o^-_{<t})}{\pi_{\mathrm{ref}}(o^-_t | q, o^-_{<t})} - \beta \log \frac{\pi_\theta(o^+_t | q, o^+_{<t})}{\pi_{\mathrm{ref}}(o^+_t | q, o^+_{<t})} \right)

    • Proximal Policy Optimization (PPO): qPsft(Q),oπθ(Oq)q \sim P_{\mathrm{sft}}(Q), o \sim \pi_\theta(O|q), GCPPO=AtGC_{\mathrm{PPO}} = A_t, where AtA_t is Generalized Advantage Estimation (GAE) computed with an auxiliary learned value network VψV_\psi.
    • Group Relative Policy Optimization (GRPO): qPsft(Q),{oi}i=1Gπθ(Oq)q \sim P_{\mathrm{sft}}(Q), \{o_i\}_{i=1}^G \sim \pi_\theta(O|q), with group-relative advantage A^i,t\hat{A}_{i,t} and exact KL regularizer gradient:

    GCGRPO(q,oi,t)=A^i,t+β(πref(oi,tq,oi,<t)πθ(oi,tq,oi,<t)1)GC_{\mathrm{GRPO}}(q, o_i, t) = \hat{A}_{i,t} + \beta \left( \frac{\pi_{\mathrm{ref}}(o_{i,t} | q, o_{i,<t})}{\pi_\theta(o_{i,t} | q, o_{i,<t})} - 1 \right)

  5. Knowl 5 — Iterative Web Data Mining Pipeline for Mathematical Text

    model/method

    The DeepSeekMath pre-training corpus (120B tokens, 35.5M web pages) is constructed from Common Crawl through an iterative data retrieval and classification pipeline:

    1. Initial Classifier Training: OpenWebMath serves as the initial positive seed. A fastText classifier is trained on 500,000 positive instances from the seed and 500,000 negative instances randomly sampled from Common Crawl (vector dimension = 256, learning rate = 0.1, word n-gram size 3\le 3, minimum word occurrences = 3, 3 training epochs).
    2. Candidate Retrieval and Scoring: The classifier scores 40B deduplicated HTML web pages from Common Crawl. Top-scoring pages are ranked, preserving the top 40B tokens in the initial iteration.
    3. Domain Identification: All pages in Common Crawl are clustered by domain (shared base URL). Domains where >10%>10\% of pages were recalled by the classifier are classified as math-related domains.
    4. Seed Expansion: Human annotators review and label specific URL paths associated with mathematical content within the identified domains (e.g., mathoverflow.net/questions). Uncollected pages under these paths are incorporated into the positive seed dataset.
    5. Iterative Retraining: The expanded seed dataset is used to retrain the fastText classifier for subsequent extraction rounds. The process is terminated after 4 iterations when new recalls drop below 2% of the cumulative dataset.
    6. Decontamination: Any web page containing an exact 10-gram match with text in downstream evaluation benchmarks (GSM8K, MATH, CMATH, AGIEval) is removed. For benchmark strings shorter than 10 tokens but 3\ge 3 tokens, exact string matching is enforced.
  6. Knowl 6 — Pre-Training Configuration and Data Mixture of DeepSeekMath-Base 7B

    experimental setup

    DeepSeekMath-Base 7B is initialized from DeepSeek-Coder-Base-v1.5 7B (using the model checkpoint prior to learning rate decay) and continually pre-trained on 500B tokens.

    The 500B-token data mixture comprises:

    • 56% DeepSeekMath Corpus (web-filtered mathematical text)
    • 20% GitHub code
    • 10% arXiv scientific papers
    • 10% Common Crawl natural language (English and Chinese)
    • 4% AlgebraicStack (mathematical code)

    Training is conducted using the HAI-LLM framework with the AdamW optimizer (β1=0.9\beta_1 = 0.9, β2=0.95\beta_2 = 0.95, weight_decay=0.1\text{weight\_decay} = 0.1). The maximum learning rate is set to 4.2×1044.2 \times 10^{-4} with a 2,000-step linear warmup. The learning rate decays to 31.6% of its peak value after 80% of total steps, and to 10.0% of peak value after 90% of total steps. The batch size is 10M tokens with a 4,096 token context length.

  7. Knowl 7 — Comparative Evaluation of DeepSeekMath Across Mathematical Reasoning Benchmarks

    data/table

    DeepSeekMath models demonstrate competitive performance across competition-level and standardized mathematical benchmarks in both English and Chinese, evaluated via few-shot chain-of-thought (CoT) prompting without external tool use:

    Model Size GSM8K MATH MMLU-STEM CMATH
    Base Models
    Minerva 7B 16.2% 14.1% 35.6% -
    Minerva 540B 58.8% 33.6% 63.9% -
    Mistral 7B 40.3% 14.3% 51.1% 44.9%
    Llemma 34B 54.0% 25.3% 52.9% 56.1%
    DeepSeekMath-Base 7B 64.2% 36.2% 56.5% 71.7%
    Instruction-Tuned RL Models
    Qwen 72B 78.9% 35.2% - -
    Math-Shepherd-Mistral 7B 84.1% 33.0% - -
    WizardMath-v1.1 7B 83.2% 33.0% - -
    GLM-4 - 87.6% 47.9% - -
    GPT-4 - 92.0% 52.9% - 86.0%
    Gemini Ultra - 94.4% 53.2% - -
    DeepSeekMath-Instruct 7B 82.9% 46.8% - 84.6%
    DeepSeekMath-RL 7B 88.2% 51.7% - 88.8%

    DeepSeekMath-Base 7B outperforms Minerva 540B on GSM8K (64.2% vs 58.8%) and MATH (36.2% vs 33.6%). DeepSeekMath-RL 7B achieves 51.7% on MATH without external tools, approaching GPT-4 (52.9%) and Gemini Ultra (53.2%), and reaches 60.9% on MATH when evaluated using self-consistency over 64 samples.

  8. Knowl 8 — Code Training Benefits for Mathematical Reasoning

    empirical result

    Pre-training language models on code data directly enhances their subsequent mathematical reasoning performance, both for self-contained text reasoning and program-aided tool use.

    Evaluating DeepSeek-LLM 1.3B under controlled training configurations demonstrates that:

    • Two-stage training with code (400B code tokens followed by 150B math tokens) yields higher chain-of-thought accuracy on GSM8K (21.9%), MATH (15.3%), and CMATH (39.7%) compared to two-stage general training (400B general tokens followed by 150B math tokens: 19.1% GSM8K, 14.4% MATH, 37.2% CMATH).
    • Code pre-training prior to math training enhances program-aided tool use (GSM8K+Python reaches 17.4% and MATH+Python reaches 9.4%, versus 14.3% and 6.7% for general pre-training followed by math).
    • One-stage mixed training (concurrently mixing 400B code and 150B math tokens) mitigates catastrophic forgetting of coding capabilities (HumanEval Pass@1 reaches 29.3% and MBPP Pass@1 reaches 39.4%, compared to 12.2% and 17.0% under two-stage sequential training) while maximizing program-aided math performance (GSM8K+Python: 19.7%, MATH+Python: 13.5%).
  9. Knowl 9 — Inefficacy of Standalone arXiv Corpora in Improving Math Reasoning Benchmarks

    empirical result

    Continual pre-training exclusively on scientific arXiv corpora fails to produce noticeable improvements on elementary and competition-level math benchmarks, and in some cases degrades downstream performance.

    Empirical comparisons across model sizes and arXiv datasets demonstrate:

    • On DeepSeek-LLM 1.3B (trained for 150B tokens), arXiv datasets (MathPile: 8.9B tokens, >85% arXiv; ArXiv-RedPajama: 28B tokens of arXiv LaTeX) achieved 2.7% and 3.3% on GSM8K, and 3.3% and 3.4% on MATH, showing negligible difference from the unadapted base model (2.9% GSM8K, 3.0% MATH).
    • On DeepSeek-Coder-Base-v1.5 7B (trained for 40B tokens), MathPile training decreased GSM8K accuracy from 29.0% to 23.6% and MATH from 12.5% to 11.5%. ArXiv-RedPajama training decreased GSM8K to 28.1% and MATH to 11.1%.
    • For formal autoformalization in Isabelle evaluated on miniF2F-test, the base model DeepSeek-Coder-Base-v1.5 7B scored 21.7%, whereas training on MathPile dropped performance to 16.4% and ArXiv-RedPajama dropped performance to 11.9%.
  10. Knowl 10 — RL Improves Majority-Voting Accuracy Without Expanding Pass@K Coverage

    empirical result

    Reinforcement learning (GRPO) on mathematical reasoning benchmarks improves majority voting performance (Maj@K) by sharpening output probability distributions around correct solutions, but does not increase generation coverage (Pass@K).

    When evaluating DeepSeekMath 7B Instruct versus DeepSeekMath-RL 7B across sample sizes K{1,4,8,16,32,64}K \in \{1, 4, 8, 16, 32, 64\} with nucleus sampling at temperature 0.7 on GSM8K and MATH:

    • Maj@K improves consistently across all KK after RL training (e.g., greedy/Maj@1 on MATH increases from 46.8% to 51.7%, and Maj@64 increases from ~57% to ~61%).
    • Pass@K curves for the Instruct model and the RL model overlap almost identically across all values of KK up to 64 (e.g., Pass@64 remains at ~87% on MATH and ~97% on GSM8K for both models).

    This indicates that RL fine-tuning mitigates post-SFT generation misalignment and boosts correct trajectory selection rather than unlocking novel reasoning primitives beyond the base/instruct model's search space.

  11. Knowl 11 — Supervised Fine-Tuning and Reinforcement Learning Setup for DeepSeekMath

    experimental setup

    The post-training alignment pipeline for DeepSeekMath-Instruct 7B and DeepSeekMath-RL 7B consists of two successive phases:

    1. Mathematical Supervised Fine-Tuning (SFT):

      • Dataset: 776,000 problem-solution pairs spanning English (GSM8K, MATH, MathInstruct, Lila-OOD) and Chinese (K-12 math across 76 topics) formatted into Chain-of-Thought (CoT), Program-of-Thought (PoT), and Tool-Integrated reasoning.
      • Optimization: Trained for 500 steps with batch size 256, constant learning rate 5×1055 \times 10^{-5}, and sequence context length 4,096 tokens.
    2. Reinforcement Learning (RL via GRPO):

      • Dataset: 144,000 chain-of-thought question prompts sampled exclusively from English GSM8K and MATH SFT data.
      • Reward Model: DeepSeekMath-Base 7B initialized with learning rate 2×1052 \times 10^{-5} trained on rule-annotated step/outcome data.
      • Policy Optimization: Initialized from DeepSeekMath-Instruct 7B, trained with policy learning rate 1×1061 \times 10^{-6}, KL coefficient β=0.04\beta = 0.04, group size G=64G = 64, maximum sequence length 1,024, prompt batch size 1,024, and a single optimization step per exploration cycle (μ=1\mu = 1).
  12. Knowl 12 — Limitations in Geometry, Formal Proving, and In-Context Learning

    limitation

    DeepSeekMath exhibits specific operational limitations relative to large closed frontier models:

    1. Geometry and Formal Theorem Proving: The model demonstrates weaker performance on geometry reasoning (such as complex planar geometry problems involving ellipses or triangles) and formal proofs compared to models like GPT-4, likely stemming from data selection biases in text-based web mining.
    2. In-Context Few-Shot Sensitivity: Unlike massive scale models where few-shot demonstrations improve benchmark accuracy, DeepSeekMath 7B shows comparable accuracy in zero-shot versus few-shot prompting, reflecting limited parameter scale in-context adaptation capabilities.

Coverage note — Omitted minor secondary evaluations on broad non-math NLP benchmarks (such as detailed multi-task subsets of MMLU and BBH) and intermediate hyperparameter sweeps for fastText embedding dimensions, as the core mathematical contributions and RL formulations are fully represented.

References

  1. 1.R. Anil, S. Borgeaud, Y. Wu, J. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, D. Silver, S. Petrov, M. Johnson, I. Antonoglou, J. Schrittwieser, A. Glaese, J. Chen, E. Pitler, T. P. Lillicrap, A. Lazaridou, O. Firat, J. Molloy, M. Isard, P. R. Barham, T. Hennigan, B. Lee, F. Viola, M. Reynolds, Y. Xu, R. Doherty, E. Collins, C. Meyer, E. Rutherford, E. Moreira, K. Ayoub, M. Goel, G. Tucker, E. Piqueras, M. Krikun, I. Barr, N. Savinov, I. Danihelka, B. Roelofs, A. White, A. Andreassen, T. von Glehn, L. Yagati, M. Kazemi, L. Gonzalez, M. Khalman, J. Sygnowski, and et al. Gemini: A family of highly capable multimodal models. CoRR, abs/2312.11805, 2023. doi: 10.48550/ARXIV.2312.11805. URL https://doi.org/10.48550/arXiv.2312.11805.
  2. 2.J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  3. 3.Z. Azerbayev, H. Schoelkopf, K. Paster, M. D. Santos, S. McAleer, A. Q. Jiang, J. Deng, S. Biderman, and S. Welleck. Llemma: An open language model for mathematics. arXiv preprint arXiv:2310.10631, 2023.
  4. 4.J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023.
  5. 5.C. Burns, P. Izmailov, J. H. Kirchner, B. Baker, L. Gao, L. Aschenbrenner, Y. Chen, A. Ecoffet, M. Joglekar, J. Leike, et al. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. arXiv preprint arXiv:2312.09390, 2023.
  6. 6.ChatGLM3 Team. Chatglm3 series: Open bilingual chat llms, 2023. URL https://github.com/THUDM/ChatGLM3.
  7. 7.M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba. Evaluating large language models trained on code. CoRR, abs/2107.03374, 2021. URL https://arxiv.org/abs/2107.03374.
  8. 8.W. Chen, X. Ma, X. Wang, and W. W. Cohen. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. CoRR, abs/2211.12588, 2022. doi: 10.48550/ARXIV.2211.12588. URL https://doi.org/10.48550/arXiv.2211.12588.
  9. 9.K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  10. 10.T. Computer. Redpajama: an open dataset for training large language models, Oct. 2023. URL https://github.com/togethercomputer/RedPajama-Data.
  11. 11.DeepSeek-AI. Deepseek LLM: scaling open-source language models with longtermism. CoRR, abs/2401.02954, 2024. doi: 10.48550/ARXIV.2401.02954. URL https://doi.org/10.48550/arXiv.2401.02954.
  12. 12.Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335, 2022.
  13. 13.L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig. PAL: program-aided language models. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, editors, International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 10764–10799. PMLR, 2023. URL https://proceedings.mlr.press/v202/gao23f.html.
  14. 14.Z. Gou, Z. Shao, Y. Gong, Y. Shen, Y. Yang, M. Huang, N. Duan, and W. Chen. Tora: A tool-integrated reasoning agent for mathematical problem solving. CoRR, abs/2309.17452, 2023. doi: 10.48550/ARXIV.2309.17452. URL https://doi.org/10.48550/arXiv.2309.17452.
  15. 15.D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. K. Li, F. Luo, Y. Xiong, and W. Liang. Deepseek-coder: When the large language model meets programming – the rise of code intelligence, 2024.
  16. 16.D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  17. 17.D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021.
  18. 18.High-flyer. Hai-llm: 高效且轻量的大模型训练工具, 2023. URL https://www.high-flyer.cn/en/blog/hai-llm.
  19. 19.Inflection AI. Inflection-2, 2023. URL https://inflection.ai/inflection-2.
  20. 20.A. Q. Jiang, S. Welleck, J. P. Zhou, W. Li, J. Liu, M. Jamnik, T. Lacroix, Y. Wu, and G. Lample. Draft, sketch, and prove: Guiding formal theorem provers with informal proofs. arXiv preprint arXiv:2210.12283, 2022.
  21. 21.A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  22. 22.A. Joulin, E. Grave, P. Bojanowski, M. Douze, H. Jégou, and T. Mikolov. Fasttext. zip: Compressing text classification models. arXiv preprint arXiv:1612.03651, 2016.
  23. 23.W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023.
  24. 24.Y. Leviathan, M. Kalman, and Y. Matias. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning, pages 19274–19286. PMLR, 2023.
  25. 25.A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, et al. Solving quantitative reasoning problems with language models. Advances in Neural Information Processing Systems, 35:3843–3857, 2022a.
  26. 26.A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y. Wu, B. Neyshabur, G. Gur-Ari, and V. Misra. Solving quantitative reasoning problems with language models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, 2022b. URL http://papers.nips.cc/paper_files/paper/2022/hash/18abbeef8cfe9203fdf9053c9c4fe191-Abstract-Conference.html.
  27. 27.H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe. Let’s verify step by step. arXiv preprint arXiv:2305.20050, 2023.
  28. 28.I. Loshchilov and F. Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  29. 29.H. Luo, Q. Sun, C. Xu, P. Zhao, J. Lou, C. Tao, X. Geng, Q. Lin, S. Chen, and D. Zhang. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. arXiv preprint arXiv:2308.09583, 2023.
  30. 30.S. Mishra, M. Finlayson, P. Lu, L. Tang, S. Welleck, C. Baral, T. Rajpurohit, O. Tafjord, A. Sabharwal, P. Clark, and A. Kalyan. LILA: A unified benchmark for mathematical reasoning. In Y. Goldberg, Z. Kozareva, and Y. Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 5807–5832. Association for Computational Linguistics, 2022. doi: 10.18653/V1/2022.EMNLP-MAIN.392. URL https://doi.org/10.18653/v1/2022.emnlp-main.392.
  31. 31.X. Nguyen, W. Zhang, X. Li, M. M. Aljunied, Q. Tan, L. Cheng, G. Chen, Y. Deng, S. Yang, C. Liu, H. Zhang, and L. Bing. Seallms - large language models for southeast asia. CoRR, abs/2312.00738, 2023. doi: 10.48550/ARXIV.2312.00738. URL https://doi.org/10.48550/arXiv.2312.00738.
  32. 32.OpenAI. GPT4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  33. 33.L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  34. 34.K. Paster, M. D. Santos, Z. Azerbayev, and J. Ba. Openwebmath: An open dataset of high-quality mathematical web text. CoRR, abs/2310.06786, 2023. doi: 10.48550/ARXIV.2310.06786. URL https://doi.org/10.48550/arXiv.2310.06786.
  35. 35.L. C. Paulson. Three years of experience with sledgehammer, a practical link between automatic and interactive theorem provers. In R. A. Schmidt, S. Schulz, and B. Konev, editors, Proceedings of the 2nd Workshop on Practical Aspects of Automated Reasoning, PAAR-2010, Edinburgh, Scotland, UK, July 14, 2010, volume 9 of EPiC Series in Computing, pages 1–10. EasyChair, 2010. doi: 10.29007/TNFD. URL https://doi.org/10.29007/tnfd.
  36. 36.S. Polu and I. Sutskever. Generative language modeling for automated theorem proving. CoRR, abs/2009.03393, 2020. URL https://arxiv.org/abs/2009.03393.
  37. 37.R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn. Direct preference optimization: Your language model is secretly a reward model. 2023.
  38. 38.J. Schulman. Approximating kl divergence, 2020. URL http://joschu.net/blog/kl-approx.html.
  39. 39.J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel. High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438, 2015.
  40. 40.J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  41. 41.F. Shi, M. Suzgun, M. Freitag, X. Wang, S. Srivats, S. Vosoughi, H. W. Chung, Y. Tay, S. Ruder, D. Zhou, D. Das, and J. Wei. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/pdf?id=fR3wGCk-IXp.
  42. 42.F. Song, B. Yu, M. Li, H. Yu, F. Huang, Y. Li, and H. Wang. Preference ranking optimization for human alignment. arXiv preprint arXiv:2306.17492, 2023.
  43. 43.M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, et al. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
  44. 44.T. Tao. Embracing change and resetting expectations, 2023. URL https://unlocked.microsoft.com/ai-anthology/terence-tao/.
  45. 45.H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. Canton-Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288, 2023. doi: 10.48550/arXiv.2307.09288. URL https://doi.org/10.48550/arXiv.2307.09288.
  46. 46.T. H. Trinh, Y. Wu, Q. V. Le, H. He, and T. Luong. Solving olympiad geometry without human demonstrations. Nature, 625(7995):476–482, 2024.
  47. 47.P. Wang, L. Li, L. Chen, F. Song, B. Lin, Y. Cao, T. Liu, and Z. Sui. Making large language models better reasoners with alignment. arXiv preprint arXiv:2309.02144, 2023a.
  48. 48.P. Wang, L. Li, Z. Shao, R. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui. Math-shepherd: Verify and reinforce llms step-by-step without human annotations. CoRR, abs/2312.08935, 2023b.
  49. 49.Z. Wang, R. Xia, and P. Liu. Generative AI for math: Part I - mathpile: A billion-token-scale pretraining corpus for math. CoRR, abs/2312.17120, 2023c. doi: 10.48550/ARXIV.2312.17120. URL https://doi.org/10.48550/arXiv.2312.17120.
  50. 50.J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS, 2022. URL http://papers.nips.cc/paper_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html.
  51. 51.T. Wei, J. Luan, W. Liu, S. Dong, and B. Wang. Cmath: Can your language model pass chinese elementary school math test?, 2023.
  52. 52.M. Wenzel, L. C. Paulson, and T. Nipkow. The isabelle framework. In O. A. Mohamed, C. A. Muñoz, and S. Tahar, editors, Theorem Proving in Higher Order Logics, 21st International Conference, TPHOLs 2008, Montreal, Canada, August 18-21, 2008. Proceedings, volume 5170 of Lecture Notes in Computer Science, pages 33–38. Springer, 2008. doi: 10.1007/978-3-540-71067-7_7. URL https://doi.org/10.1007/978-3-540-71067-7_7.
  53. 53.H. Xia, T. Ge, P. Wang, S.-Q. Chen, F. Wei, and Z. Sui. Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation. In H. Bouamor, J. Pino, and K. Bali, editors, Findings of the Association for Computational Linguistics: EMNLP 2023, pages 3909–3925, Singapore, Dec. 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-emnlp.257. URL https://aclanthology.org/2023.findings-emnlp.257.
  54. 54.H. Xia, Z. Yang, Q. Dong, P. Wang, Y. Li, T. Ge, T. Liu, W. Li, and Z. Sui. Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding. arXiv preprint arXiv:2401.07851, 2024.
  55. 55.S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601, 2023.
  56. 56.L. Yu, W. Jiang, H. Shi, J. Yu, Z. Liu, Y. Zhang, J. T. Kwok, Z. Li, A. Weller, and W. Liu. Metamath: Bootstrap your own mathematical questions for large language models. CoRR, abs/2309.12284, 2023. doi: 10.48550/ARXIV.2309.12284. URL https://doi.org/10.48550/arXiv.2309.12284.
  57. 57.Z. Yuan, H. Yuan, C. Li, G. Dong, C. Tan, and C. Zhou. Scaling relationship on learning mathematical reasoning with large language models. arXiv preprint arXiv:2308.01825, 2023a.
  58. 58.Z. Yuan, H. Yuan, C. Tan, W. Wang, S. Huang, and F. Huang. Rrhf: Rank responses to align language models with human feedback without tears. arXiv preprint arXiv:2304.05302, 2023b.
  59. 59.X. Yue, X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su, and W. Chen. Mammoth: Building math generalist models through hybrid instruction tuning. CoRR, abs/2309.05653, 2023. doi: 10.48550/ARXIV.2309.05653. URL https://doi.org/10.48550/arXiv.2309.05653.
  60. 60.K. Zheng, J. M. Han, and S. Polu. Minif2f: a cross-system benchmark for formal olympiad-level mathematics. arXiv preprint arXiv:2109.00110, 2021.
  61. 61.W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan. AGIEval: A human-centric benchmark for evaluating foundation models. CoRR, abs/2304.06364, 2023. doi: 10.48550/arXiv.2304.06364. URL https://doi.org/10.48550/arXiv.2304.06364.

Citation

MLA
Shao, Z., et al. “DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models”. arXiv, 2024, http://arxiv.org/abs/2402.03300v3.
APA
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y. K., Wu, Y., & Guo, D. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv. http://arxiv.org/abs/2402.03300v3
Chicago
Shao, Z., P. Wang, Q. Zhu, et al. 2024. “DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models”. arXiv. http://arxiv.org/abs/2402.03300v3.
Harvard
Shao, Z. et al. (2024) “DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.03300v3.
Vancouver
1. Shao Z, Wang P, Zhu Q, et al (2024) DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv

BibTeX

@article{shao2024deepseekmath,
  title = {DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models},
  author = {Shao, Zhihong and Wang, Peiyi and Zhu, Qihao and Xu, Runxin and Song, Junxiao and Bi, Xiao and Zhang, Haowei and Zhang, Mingchuan and Li, Y. K. and Wu, Y. and Guo, Daya},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.03300v3},
  eprint = {2402.03300}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors