RAGEN-2: Reasoning Collapse in Agentic RL

Zihan WangChi GuiXing JinQineng WangLicheng LiuKangrui WangShiqi ChenLinjie LiZhengyuan YangPingyue Zhang

article2026arXiv14 citations

Reveals that reinforcement learning agents often succumb to input-agnostic template collapse undetected by standard entropy metrics, and introduces mutual information diagnostics alongside signal-to-noise prompt filtering to restore responsive reasoning and improve performance.

Listen

Training multi-turn artificial intelligence agents using reinforcement learning is notoriously unstable. Practitioners commonly monitor process stability using entropy, which measures output diversity. However, the article reveals that entropy only tracks diversity within the same input, failing to detect whether the model's reasoning actually adapts across different tasks and prompts. Consequently, models can experience "template collapse," where they produce superficially diverse yet completely generic, input-agnostic reasoning templates. Because this failure mode remains invisible to standard entropy and reward metrics, models can silently degrade into unreliable decision-makers.

The main objective of the article is to diagnose template collapse, explain its underlying mathematical causes, and demonstrate an effective, computationally lightweight mitigation strategy. The researchers formulate an information-theoretic framework to measure true input dependence and test an adaptive optimization method to maintain high reasoning quality.

To accomplish this, the authors evaluate large language model agents across seven complementary synthetic testbeds spanning irreversible spatial planning, stochastic grid navigation, symbolic math reasoning, web search, online shopping navigation, and competitive code generation. They analyze multiple reinforcement learning algorithms—including standard policy optimization and group-relative variants—across diverse model architectures ranging from 0.5 billion to 7 billion parameters, as well as multimodal vision-language models. To track input sensitivity online without needing external evaluation models, the authors introduce a mutual information proxy computed via in-batch cross-scoring.

The article establishes several key findings. First, mutual information proxies reliably diagnose reasoning degradation early in training, demonstrating a positive correlation with final task performance (+0.39), whereas standard entropy metrics correlate negatively (−0.11 to −0.14) and point in the wrong direction. Second, the authors explain template collapse through a signal-to-noise ratio mechanism: when within-prompt reward variance is near zero, task-discriminative gradients vanish, allowing input-agnostic regularizers (such as entropy bonuses and reference-model constraints) to dominate and erase input-specific reasoning. Third, the proposed intervention—SNR-Aware Filtering, which dynamically retains only high-reward-variance prompts before computing parameter updates—consistently improves peak performance across tasks, models, and modalities. For example, in planning and navigation benchmarks, filtering increased task success rates by up to 16 to 59 percentage points while simultaneously cutting per-step gradient computation time by 26% to 41%.

These findings imply that conventional training pipelines often waste computational resources on uninformative data that actively harms reasoning capabilities. Standard stabilization techniques, such as adjusting entropy or penalty coefficients, cannot prevent template collapse because they do not address the underlying signal-to-noise imbalance. Implementing variance-based filtering directly enhances gradient signal quality, providing a practical, cost-effective way to improve agent robustness and learning efficiency without requiring extra models or supervision.

Based on these results, engineering teams should replace or supplement entropy tracking with mutual information proxies to monitor online agent training. Training pipelines should adopt adaptive, nucleus-style reward variance filtering to eliminate low-signal updates. Practitioners can run a quick diagnostic check—measuring the dispersion ratio of reward variance across a single rollout batch—to determine beforehand whether an environment has sufficient variance heterogeneity to benefit from filtering.

While the findings are robust across diverse single-agent tasks and modalities, certain limitations apply. Reward variance acts as an effective signal proxy only when environments possess sufficient variance heterogeneity; in purely noisy or uniformly sparse environments, variance filtering offers diminished benefits. Furthermore, the framework has not yet been validated in multi-agent settings, and extremely aggressive filtering parameters could potentially narrow exploration if not tuned per task. Readers should view the method as a proven stabilizer for single-agent optimization that requires task-specific calibration of data retention rates.

arXiv: 2604.06268mll-lab-nu/RAGEN
Cover for RAGEN-2: Reasoning Collapse in Agentic RL

Abstract

RL training of multi-turn LLM agents is inherently unstable, and reasoning quality directly determines task performance. Entropy is widely used to track reasoning stability. However, entropy only measures diversity within the same input, and cannot tell whether reasoning actually responds to different inputs. In RAGEN-2, we find that even with stable entropy, models can rely on fixed templates that look diverse but are input-agnostic. We call this template collapse, a failure mode invisible to entropy and all existing metrics. To diagnose this failure, we decompose reasoning quality into within-input diversity (Entropy) and cross-input distinguishability (Mutual Information, MI), and introduce a family of mutual information proxies for online diagnosis. Across diverse tasks, mutual information correlates with final performance much more strongly than entropy, making it a more reliable proxy for reasoning quality. We further explain template collapse with a signal-to-noise ratio (SNR) mechanism. Low reward variance weakens task gradients, letting regularization terms dominate and erase cross-input reasoning differences. To address this, we propose SNR-Aware Filtering to select high-signal prompts per iteration using reward variance as a lightweight proxy. Across planning, math reasoning, web navigation, and code execution, the method consistently improves both input dependence and task performance.

Table of Contents

  • 1 Introduction
  • 2 Template Collapse in Multi-turn Agent RL
  • 2.1 Setup and Preliminaries
  • 2.2 Rethinking Reasoning Collapse from an Information-Theoretic Lens
  • 2.3 Mutual Information Proxy Family
  • 3 The Mechanism of Template Collapse: A Signal-to-Noise Ratio (SNR) View
  • 3.1 Observing Signal-Noise Imbalance in RL Gradients
  • 3.2 Formalizing the SNR Mechanism via Gradient Decomposition
  • 3.3 SNR-Aware Filtering: Prioritizing High-Signal Updates
  • 4 Experiments
  • 4.1 Experimental Testbed
  • 4.2 Template Collapse as a Consistent Failure Mode
  • 4.3 SNR-Aware Filtering Consistently Improves Performance
  • 5 Analysis
  • 5.1 MI Diagnoses Collapse Better Than Entropy Across All Interventions
  • 5.2 Does SNR Mechanism Really Interpret Agent RL?
  • 6 Related Work
  • 7 Conclusions and Limitations
  • 8 Acknowledgements
  • References
  • A Extended Related Work
  • B Detailed Experimental Settings
  • B.1 Environments and Tasks
  • B.2 Training and Evaluation Setup
  • C Filtering Ablation Results
  • D Additional Experimental Visualizations
  • D.1 MI Proxy Metrics During Training
  • E Notation and basic identities
  • E.1 Random variables and distributions
  • E.2 Entropy and mutual information
  • F Scorer-based Proxies for Reasoning Diversity
  • F.1 Setup and notation
  • G Formal Definition of the Filtering Operator
  • G.1 Filtering Strategy Variants
  • H RV Controls Task-Signal Magnitude and SNR
  • H.1 Setup
  • H.2 Assumption
  • H.3 Task-gradient magnitude is RV-controlled
  • H.4 SNR is upper bounded by RV and reward noise
  • H.5 Low-SNR updates induce parameter drift
  • I Template Mixing Reduces Input Dependence
  • J Filtering Reduces Gradient-Estimation MSE
  • J.1 Setup
  • J.2 Unfiltered vs. filtered estimators
  • K Reward-Agnostic Regularizers and Update Dominance
  • K.1 Setup
  • K.2 Low-RV prompts amplify regularizer influence
  • L KL-Closeness to the Base Implies MI-Closeness
  • M Decomposing Changes in Input Dependence
  • N GRPO Normalization Amplifies Noise at Low RV
  • O Core Author Contributions

Knowls

  1. Knowl 1 — Template Collapse in Agentic Reinforcement Learning

    definition

    In closed-loop multi-turn agent reinforcement learning, where an agent policy πθ\pi_\theta receives input context XX (comprising prompts, prior observations, actions, and previous reasoning) and generates reasoning token sequence ZZ, marginal reasoning diversity H(Z)H(Z) decomposes into:

    H(Z)=I(X;Z)+H(Z∣X)H(Z) = I(X; Z) + H(Z \mid X)

    where I(X;Z)=EX∼P(X),Z∼πθ(⋅∣X)[log⁡πθ(Z∣X)pθ(Z)]I(X; Z) = \mathbb{E}_{X \sim P(X), Z \sim \pi_\theta(\cdot \mid X)} \left[ \log \frac{\pi_\theta(Z \mid X)}{p_\theta(Z)} \right] measures cross-input dependence (how much the reasoning reflects the specific input prompt), and H(Z∣X)=−EX∼P(X),Z∼πθ(⋅∣X)[log⁡πθ(Z∣X)]H(Z \mid X) = -\mathbb{E}_{X \sim P(X), Z \sim \pi_\theta(\cdot \mid X)} [\log \pi_\theta(Z \mid X)] measures within-input diversity (conditional entropy).

    Reasoning outputs fall into four distinct operational regimes:

    1. Diverse Reasoning (high H(Z∣X)H(Z \mid X), high I(X;Z)I(X; Z)): Reasoning varies across rollouts for the same input while remaining systematically grounded and distinguishable across different inputs.
    2. Template Collapse (high H(Z∣X)H(Z \mid X), low I(X;Z)I(X; Z)): Reasoning appears superficially diverse within single inputs but degenerates into input-agnostic templates across different inputs. Existing entropy metrics fail to detect this failure because H(Z∣X)H(Z \mid X) remains stable or elevated.
    3. Compressed Reasoning (low H(Z∣X)H(Z \mid X), high I(X;Z)I(X; Z)): Reasoning is faithful to the input prompt but overly deterministic across rollouts.
    4. Low-Entropy Collapse (low H(Z∣X)H(Z \mid X), low I(X;Z)I(X; Z)): Reasoning is fully degenerate, exhibiting both deterministic and input-agnostic outputs.
  2. Knowl 2 — In-Batch Cross-Scoring Proxies for Mutual Information

    model/method

    To quantify input dependence I(X;Z)I(X; Z) without external models during training rollouts of PP prompts with GG reasoning traces per prompt, the policy πθ\pi_\theta computes a teacher-forced scoring matrix Li,k,j=log⁡pθ(Zi,k∣Xj)L_{i,k,j} = \log p_\theta(Z_{i,k} \mid X_j), where Zi,kZ_{i,k} is the kk-th rollout generated for prompt XiX_i and evaluated under prompt XjX_j. Two length-normalized quantities are computed:

    matchedi,k=Li,k,i∣Zi,k∣,marginali,k=1∣Zi,k∣log⁡(1P∑j=1Pexp⁡(Li,k,j))\text{matched}_{i,k} = \frac{L_{i,k,i}}{|Z_{i,k}|}, \quad \text{marginal}_{i,k} = \frac{1}{|Z_{i,k}|} \log \left( \frac{1}{P} \sum_{j=1}^P \exp(L_{i,k,j}) \right)

    where matchedi,k\text{matched}_{i,k} is the per-token log-likelihood of Zi,kZ_{i,k} under its true source prompt XiX_i, and marginali,k\text{marginal}_{i,k} estimates the marginal log-likelihood under a uniform empirical prompt mixture.

    Two primary estimators are derived:

    1. Retrieval Accuracy (Retrieval-Acc): A discrete metric defined as: Acc=1PG∑i=1P∑k=1GI[i=arg⁡max⁡jLi,k,j]\text{Acc} = \frac{1}{P G} \sum_{i=1}^P \sum_{k=1}^G \mathbb{I}\left[ i = \arg\max_j L_{i,k,j} \right] Under template collapse, Acc\text{Acc} approaches random chance level 1/P1/P.

    2. MI-ZScore-EMA: A continuous mutual information estimate I^(X;Z)=1PG∑i=1P∑k=1G(matchedi,k−marginali,k)\widehat{I}(X; Z) = \frac{1}{P G} \sum_{i=1}^P \sum_{k=1}^G (\text{matched}_{i,k} - \text{marginal}_{i,k}), normalized by the batch marginal standard deviation σbatch\sigma_{\text{batch}} and stabilized via exponential moving average: MI-ZScore-EMA=1PG∑i=1P∑k=1Gmatchedi,k−marginali,kσEMA+ϵ\text{MI-ZScore-EMA} = \frac{1}{P G} \sum_{i=1}^P \sum_{k=1}^G \frac{\text{matched}_{i,k} - \text{marginal}_{i,k}}{\sigma_{\text{EMA}} + \epsilon} where σEMA(t)=ασEMA(t−1)+(1−α)σbatch(t)\sigma_{\text{EMA}}^{(t)} = \alpha \sigma_{\text{EMA}}^{(t-1)} + (1 - \alpha) \sigma_{\text{batch}}^{(t)} with smoothing factor α=0.9\alpha = 0.9 and ϵ=10−3\epsilon = 10^{-3}.

  3. Knowl 3 — Top-p SNR-Aware Filtering Algorithm

    algorithm

    SNR-Aware Filtering selects high-signal prompts per rollout iteration using sample reward variance as a lightweight proxy for signal-to-noise ratio, updating policy parameters only on prompts with sufficient reward variance.

    Input: Policy πθ\pi_\theta, batch of PP prompts {x1,…,xP}\{x_1, \dots, x_P\}, trajectories per prompt GG, keep rate parameter ρ∈(0,1]\rho \in (0, 1]
    Output: Filtered policy parameter update Δθ\Delta \theta
    for each prompt xix_i (i∈{1,…,P}i \in \{1, \dots, P\}) do
        Sample GG rollouts {zi,1,…,zi,G}∼πθ(⋅∣xi)\{z_{i,1}, \dots, z_{i,G}\} \sim \pi_\theta(\cdot \mid x_i)
        Collect episode returns {Ri,1,…,Ri,G}\{R_{i,1}, \dots, R_{i,G}\} where Ri,g=R(zi,g;xi)R_{i,g} = R(z_{i,g}; x_i)
        Compute conditional mean return Rˉ(xi)←1G∑g=1GRi,g\bar{R}(x_i) \leftarrow \frac{1}{G} \sum_{g=1}^G R_{i,g}
        Compute sample reward variance Var^(R∣X=xi)←1G−1∑g=1G(Ri,g−Rˉ(xi))2\widehat{\text{Var}}(R \mid X = x_i) \leftarrow \frac{1}{G-1} \sum_{g=1}^G (R_{i,g} - \bar{R}(x_i))^2
    end for
    Find permutation σ\sigma of {1,…,P}\{1, \dots, P\} such that Var^(R∣X=xσ(1))≥⋯≥Var^(R∣X=xσ(P))\widehat{\text{Var}}(R \mid X = x_{\sigma(1)}) \ge \dots \ge \widehat{\text{Var}}(R \mid X = x_{\sigma(P)})
    Compute variance threshold τ←ρ∑i=1PVar^(R∣X=xi)\tau \leftarrow \rho \sum_{i=1}^P \widehat{\text{Var}}(R \mid X = x_i)
    Find cutoff index k∗←min⁡{k:∑j=1kVar^(R∣X=xσ(j))≥τ}k^* \leftarrow \min \left\{ k : \sum_{j=1}^k \widehat{\text{Var}}(R \mid X = x_{\sigma(j)}) \ge \tau \right\}
    Define selected prompt subset S←{σ(1),…,σ(k∗)}S \leftarrow \{\sigma(1), \dots, \sigma(k^*)\}
    Form filtered objective Lρ(θ)←1k∗∑i∈S∑j∈BiLθ(ξj)\mathcal{L}_\rho(\theta) \leftarrow \frac{1}{k^*} \sum_{i \in S} \sum_{j \in \mathcal{B}_i} L_\theta(\xi_j) where Bi\mathcal{B}_i is the rollout sample group for prompt xix_i
    Update policy parameters θ\theta via ∇θLρ(θ)\nabla_\theta \mathcal{L}_\rho(\theta)
  4. Knowl 4 — Upper Bound on Task-Gradient Magnitude by Within-Prompt Reward Variance

    theoretical result

    Let XX denote an input prompt, z∼πθ(⋅∣X)z \sim \pi_\theta(\cdot \mid X) a rollout trajectory, s(z;x)=∇θlog⁡πθ(z∣x)s(z; x) = \nabla_\theta \log \pi_\theta(z \mid x) the score function, b(x)=E[R(z;x)∣X=x]b(x) = \mathbb{E}[R(z; x) \mid X = x] the conditional-mean baseline, and A(z;x)=R(z;x)−b(x)A(z; x) = R(z; x) - b(x) the advantage function. Defining the reward variance as RV(x)=Var(R(z;x)∣X=x)=E[A(z;x)2∣X=x]\text{RV}(x) = \text{Var}(R(z; x) \mid X = x) = \mathbb{E}[A(z; x)^2 \mid X = x] and the task gradient as gtask(x)=E[A(z;x)s(z;x)∣X=x]g_{\text{task}}(x) = \mathbb{E}[A(z; x) s(z; x) \mid X = x], the Euclidean norm of the task gradient satisfies:

    ∥gtask(x)∥≤RV(x)E[∥s(z;x)∥2∣X=x]\|g_{\text{task}}(x)\| \le \sqrt{\text{RV}(x)} \sqrt{\mathbb{E}\left[ \|s(z; x)\|^2 \mid X = x \right]}

    When within-prompt reward variance RV(x)\text{RV}(x) approaches zero, the task-discriminative gradient gtask(x)g_{\text{task}}(x) provably vanishes, leaving updates to be dominated by regularizers.

  5. Knowl 5 — Upper Bound on Policy Gradient Signal-to-Noise Ratio under Reward Noise

    theoretical result

    Assume the observed return decomposes as R(z;x)=μ(x,z)+εR(z; x) = \mu(x, z) + \varepsilon, where μ(x,z)=E[R(z;x)∣x,z]\mu(x, z) = \mathbb{E}[R(z; x) \mid x, z] is the trajectory-dependent expected reward and ε\varepsilon is zero-mean noise with E[ε∣x,z]=0\mathbb{E}[\varepsilon \mid x, z] = 0 and Var(ε∣x,z)=σ2(x)≥0\text{Var}(\varepsilon \mid x, z) = \sigma^2(x) \ge 0. Let g^task(x)=1G∑k=1GA(zk;x)s(zk;x)\widehat{g}_{\text{task}}(x) = \frac{1}{G} \sum_{k=1}^G A(z_k; x) s(z_k; x) be the GG-sample Monte Carlo task-gradient estimator using conditional baseline b(x)=E[R∣X=x]b(x) = \mathbb{E}[R \mid X = x]. Defining the prompt-level signal-to-noise ratio as:

    SNR(x):=∥gtask(x)∥E[∥g^task(x)−gtask(x)∥2∣X=x]\text{SNR}(x) := \frac{\|g_{\text{task}}(x)\|}{\sqrt{\mathbb{E}\left[ \|\widehat{g}_{\text{task}}(x) - g_{\text{task}}(x)\|^2 \mid X = x \right]}}

    the signal-to-noise ratio is upper bounded by:

    SNR(x)≤G⋅RV(x)σ(x)\text{SNR}(x) \le \sqrt{G} \cdot \frac{\sqrt{\text{RV}(x)}}{\sigma(x)}

    When within-prompt reward variance RV(x)\text{RV}(x) is low relative to the noise level σ(x)\sigma(x), the gradient update estimator is dominated by noise.

  6. Knowl 6 — Noise Amplification Floor in GRPO Normalization at Low Reward Variance

    theoretical result

    In Group Relative Policy Optimization (GRPO), the advantage is normalized by the empirical within-prompt standard deviation: A~(z;x)=R(z;x)−b(x)RV(x)\widetilde{A}(z; x) = \frac{R(z; x) - b(x)}{\sqrt{\text{RV}(x)}}, where b(x)=E[R(z;x)∣X=x]b(x) = \mathbb{E}[R(z; x) \mid X = x] and RV(x)=Var(R∣X=x)\text{RV}(x) = \text{Var}(R \mid X = x). For KK i.i.d. rollouts z1,…,zK∼πθ(⋅∣x)z_1, \dots, z_K \sim \pi_\theta(\cdot \mid x) and score functions sk=∇θlog⁡πθ(zk∣x)s_k = \nabla_\theta \log \pi_\theta(z_k \mid x), the GRPO estimator g^GRPO(x)=1K∑k=1KA~ksk\widehat{g}_{\text{GRPO}}(x) = \frac{1}{K} \sum_{k=1}^K \widetilde{A}_k s_k has a mean-squared error floor satisfying:

    E[∥g^GRPO(x)−gGRPO(x)∥2∣X=x]≥1K⋅σ2(x)RV(x)E[∥s(z;x)∥2∣X=x]\mathbb{E}\left[ \|\widehat{g}_{\text{GRPO}}(x) - g_{\text{GRPO}}(x)\|^2 \mid X = x \right] \ge \frac{1}{K} \cdot \frac{\sigma^2(x)}{\text{RV}(x)} \mathbb{E}\left[ \|s(z; x)\|^2 \mid X = x \right]

    where σ2(x)=Var(R(z;x)∣x,z)\sigma^2(x) = \text{Var}(R(z; x) \mid x, z) is the reward noise variance. Normalizing advantages by RV(x)\sqrt{\text{RV}(x)} causes the estimator variance floor to scale as RV(x)−1\text{RV}(x)^{-1}, amplifying gradient noise on low-variance prompts.

  7. Knowl 7 — Mutual Information Contraction Under Prompt-Independent Template Mixing

    theoretical result

    Let X∼P(X)X \sim P(X) and Z∣X=x∼p(z∣x)Z \mid X = x \sim p(z \mid x). Let q(z)q(z) be any prompt-independent distribution representing a fixed reasoning template. For mixing parameter α∈[0,1]\alpha \in [0, 1], define the contaminated conditional distribution pα(z∣x)=(1−α)p(z∣x)+αq(z)p_\alpha(z \mid x) = (1 - \alpha) p(z \mid x) + \alpha q(z) and corresponding marginal pα(z)=(1−α)p(z)+αq(z)p_\alpha(z) = (1 - \alpha) p(z) + \alpha q(z).

    The resulting mutual information Iα(X;Z)I_\alpha(X; Z) satisfies:

    Iα(X;Z)≤(1−α)I(X;Z)I_\alpha(X; Z) \le (1 - \alpha) I(X; Z)

    Mixing a prompt-agnostic template into the policy output contracts cross-input mutual information by at least a factor of (1−α)(1 - \alpha).

  8. Knowl 8 — Regularizer Dominance Ratio Under Vanishing Reward Variance

    theoretical result

    Let the total expected policy update be gtotal(x)=gtask(x)+greg(x)g_{\text{total}}(x) = g_{\text{task}}(x) + g_{\text{reg}}(x), where gtask(x)=E[(R(z;x)−b(x))s(z;x)∣X=x]g_{\text{task}}(x) = \mathbb{E}[(R(z; x) - b(x)) s(z; x) \mid X = x] is the task gradient and greg(x)=λKLgKL(x)+λentgent(x)g_{\text{reg}}(x) = \lambda_{\text{KL}} g_{\text{KL}}(x) + \lambda_{\text{ent}} g_{\text{ent}}(x) represents input-agnostic regularization terms (e.g., KL divergence penalty to a reference policy and entropy bonus). Defining the regularizer dominance ratio as:

    ρ(x):=∥greg(x)∥∥gtask(x)∥+∥greg(x)∥∈[0,1]\rho(x) := \frac{\|g_{\text{reg}}(x)\|}{\|g_{\text{task}}(x)\| + \|g_{\text{reg}}(x)\|} \in [0, 1]

    it satisfies the lower bound:

    ρ(x)≥∥greg(x)∥∥greg(x)∥+RV(x)E[∥s(z;x)∥2∣X=x]\rho(x) \ge \frac{\|g_{\text{reg}}(x)\|}{\|g_{\text{reg}}(x)\| + \sqrt{\text{RV}(x)} \sqrt{\mathbb{E}[\|s(z; x)\|^2 \mid X = x]}}

    Because regularizer gradients greg(x)g_{\text{reg}}(x) remain flat across prompts while task gradients scale with RV(x)\sqrt{\text{RV}(x)}, low reward variance prompts cause ρ(x)→1\rho(x) \to 1, meaning updates on these prompts are almost entirely driven by input-agnostic regularization.

  9. Knowl 9 — Empirical Performance of SNR-Aware Filtering Across Agent RL Benchmarks

    data/table

    Across multiple RL algorithms (PPO, DAPO, GRPO, Dr. GRPO), model scales (Qwen2.5 0.5B to 7B), base model types (Qwen2.5-3B-Instruct, Llama3.2-3B), and modalities (text and vision on Qwen2.5-VL-3B), SNR-Aware Filtering with keep rate ρ=0.9\rho = 0.9 consistently improves mean task success rate over the unfiltered baseline.

    Experiment Variants Sokoban FrozenLake MetaMathQA Countdown Average
    Baseline
    PPO, Qwen2.5-3B 12.9 (+16.0) 67.0 (+10.9) 92.6 (+0.6) 97.9 (+0.0) 67.6 (+6.9)
    Algorithm
    DAPO 16.2 (+5.1) 66.8 (+2.1) 90.8 (+2.8) 95.7 (+1.6) 67.4 (+2.9)
    GRPO 12.1 (+9.0) 70.9 (-3.0) 91.2 (+1.2) 95.7 (+2.2) 67.5 (+3.7)
    Dr. GRPO 12.1 (-0.4) 23.2 (+0.6) 91.2 (+1.4) 96.5 (+1.4) 55.8 (+0.8)
    Model Scale (PPO)
    Qwen2.5-0.5B 3.3 (+22.9) 19.5 (+0.0) 10.0 (-0.2) 23.0 (-0.7) 14.0 (+5.5)
    Qwen2.5-1.5B 17.0 (+6.2) 36.5 (+1.6) 80.3 (+7.0) 56.6 (+1.6) 47.6 (+4.1)
    Qwen2.5-7B 42.4 (+4.9) 85.0 (-0.6) 84.0 (+11.7) 97.7 (+0.3) 77.3 (+4.1)
    Model Type
    Qwen2.5-3B-Instruct 22.5 (+14.2) 83.6 (+2.3) 91.2 (+0.4) 96.3 (-0.6) 73.4 (+4.1)
    Llama3.2-3B 24.4 (+18.8) 84.6 (-0.2) 86.1 (+3.7) 99.2 (-1.2) 73.6 (+5.3)
    Modality (Input Type)
    Qwen2.5-VL-3B (Text) 53.0 (+6.0) 16.0 (+53.5) - - 34.5 (+29.8)
    Qwen2.5-VL-3B (Vision) 65.0 (+12.0) 19.5 (+59.5) - - 42.3 (+35.8)

    Each entry presents the baseline peak validation success rate (%) with the delta (+Δ+\Delta) achieved by applying SNR-Aware Filtering in parentheses. Computing sample reward variance adds <0.1%<0.1\% wall-clock iteration overhead, while filtering uninformative rollouts reduces per-step backward pass time by 26%–41%26\%\text{--}41\% without increasing total rollout compute budget.

  10. Knowl 10 — Pre-Training Prediction of Filtering Efficacy via Reward-Variance Heterogeneity

    empirical result

    The effectiveness of SNR-Aware Filtering is governed by the heterogeneity of within-prompt reward variance across the prompt batch, measured by the coefficient of variation ratio Std(RV)/Mean(RV)\text{Std}(\text{RV}) / \text{Mean}(\text{RV}) computed from a single rollout batch prior to full training:

    1. In Sokoban (14B model, Std/Mean=1.29\text{Std}/\text{Mean} = 1.29, RV Mean=2.24\text{RV Mean} = 2.24, RV Std=2.88\text{RV Std} = 2.88), filtering improves performance by +4.6%+4.6\%.
    2. In Sokoban (3B model, Std/Mean=1.16\text{Std}/\text{Mean} = 1.16, RV Mean=2.49\text{RV Mean} = 2.49, RV Std=2.89\text{RV Std} = 2.89), filtering improves performance by +3.2%+3.2\%.
    3. In FrozenLake with GRPO (Std/Mean=0.33\text{Std}/\text{Mean} = 0.33, RV Mean=0.54\text{RV Mean} = 0.54, RV Std=0.18\text{RV Std} = 0.18), filtering changes performance by −5.0%-5.0\%.

    A high ratio indicates a bimodal reward-variance distribution where SNR-Aware Filtering cleanly separates informative prompts from uninformative ones. A low ratio indicates uniform reward variance across prompts, where filtering discards prompts indiscriminately.

  11. Knowl 11 — Diagnostic Superiority of Mutual Information Over Reasoning Entropy

    empirical result

    During reinforcement learning of multi-turn reasoning agents, mutual information proxies correlate positively with final task performance, whereas entropy-based metrics correlate negatively:

    • Trajectory MI-ZScore achieves a Spearman rank correlation of +0.39+0.39 with final task success.
    • MI-Seq Estimate achieves a Spearman rank correlation of +0.22+0.22.
    • Marginal Reasoning Entropy achieves a negative Spearman correlation of −0.11-0.11.
    • Conditional Entropy H(Z∣X)H(Z \mid X) achieves a negative Spearman correlation of −0.14-0.14.

    Furthermore, in a controlled quartile ablation where prompts are sorted into four reward-variance tiers (Q1 highest RV [4.4–5.6][4.4\text{--}5.6] to Q4 lowest RV [0.0–0.1][0.0\text{--}0.1]) on Sokoban with Qwen2.5-3B, models trained exclusively on Q1 achieve 21.1%21.1\% task performance and 0.950.95 MI proxy, whereas models trained on Q4 degrade monotonically to 11.0%11.0\% task performance and 0.730.73 MI proxy, confirming that higher reward variance causally drives both input dependence and task performance.

  12. Knowl 12 — Limitations of Reward-Variance Proxies and SNR-Aware Filtering

    limitation

    The methodology and theoretical framework of SNR-Aware Filtering carry several stated limitations:

    1. Gradient Coupling: The SNR decomposition assumes that task signal and regularization noise separate cleanly, but they may couple in practice through gradient accumulation.
    2. Multi-Agent Generalization: Evaluations are conducted exclusively in single-agent environments; the propagation of template collapse in multi-agent RL remains uncharacterized.
    3. Criterion Gaming: Highly capable models could potentially game the filtering criterion by artificially inflating rollout reward variance without producing meaningful task-directed reasoning.
    4. Sparse and Noisy Environments: Reward variance functions as an effective SNR proxy only when reward signals are informative; in environments with extreme stochasticity or high evaluation noise, reward variance is dominated by environmental randomness rather than policy signal.
    5. Exploration Coverage: Aggressive filtering (very low keep rate ρ\rho) risks discarding hard or edge-case prompts, requiring hyperparameter tuning of the keep mass per task.

Coverage note — None was omitted; all key theoretical definitions, information-theoretic formulations, gradient bounds, algorithms, empirical benchmark tables, quartile and ablation analyses, and limitations have been extracted into self-contained knowls.

References

  1. 1.Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2025.
  2. 2.Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. Openai gym, 2016.
  3. 3.Shiqi Chen, Tongyao Zhu, Zian Wang, Jinghan Zhang, Kangrui Wang, Siyang Gao, Teng Xiao, Yee Whye Teh, Junxian He, and Manling Li. Internalizing world models via self-play finetuning for agentic rl, 2025.
  4. 4.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems, 2021.
  5. 5.Thomas M. Cover and Joy A. Thomas. Elements of Information Theory. Wiley-Interscience, 2 edition, 2006.
  6. 6.Ganqu Cui, Lifan Yuan, Zefan Wang, Hanbin Wang, Yuchen Zhang, Jiacheng Chen, Wendi Li, Bingxiang He, Yuchen Fan, Tianyu Yu, Qixin Xu, Weize Chen, Jiarui Yuan, Huayu Chen, Kaiyan Zhang, Xingtai Lv, Shuo Wang, Yuan Yao, Xu Han, Hao Peng, Yu Cheng, Zhiyuan Liu, Maosong Sun, Bowen Zhou, and Ning Ding. Process reinforcement through implicit rewards, 2025.
  7. 7.DeepSeek AI. Deepseek-r1 incentivizes reasoning in llms through reinforcement learning. Nature, 645(8081):633–638, September 2025.
  8. 8.Zi-Yi Dou, Cheng-Fu Yang, Xueqing Wu, Kai-Wei Chang, and Nanyun Peng. Re-rest: Reflection-reinforced self-training for language agents, 2025.
  9. 9.Lang Feng, Zhenghai Xue, Tingcong Liu, and Bo An. Group-in-group policy optimization for llm agent training, 2025.
  10. 10.Lirong Gao, Ru Peng, Yiming Zhang, and Junbo Zhao. Dory: Deliberative prompt recovery for llm, 2024.
  11. 11.Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hanna Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. Evaluating models’ local decision boundaries via contrast sets, 2020.
  12. 12.Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, Daniel A. Roberts, Diyi Yang, David L. Donoho, and Sanmi Koyejo. Is model collapse inevitable? breaking the curse of recursion by accumulating real and synthetic data, 2024.
  13. 13.Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine. Soft actor-critic algorithms and applications, 2019.
  14. 14.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration, 2020.
  15. 15.Naman Jain, Kimin Han, Alex Gu, Wen-Ding Li, Feng Yan, Tianjun Zhang, Yizhou Wang, Koushik Sen, Ion Stoica, and Joseph E. Gonzalez. Livecodebench: Holistic and contamination-free evaluation of large language models for code. arXiv preprint arXiv:2403.07974, 2024.
  16. 16.Michael Katz, Harsha Kokel, and Sarath Sreedharan. Benchmarking llms on the game of countdown, 2025.
  17. 17.Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. Understanding the effects of rlhf on llm generalisation and diversity, 2024.
  18. 18.Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamilė Lukošiūtė, Karina Nguyen, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Robin Larson, Sam McCandlish, Sandipan Kundu, Saurav Kadavath, Shannon Yang, Thomas Henighan, Timothy Maxwell, Timothy Telleen-Lawton, Tristan Hume, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, and Ethan Perez. Measuring faithfulness in chain-of-thought reasoning, 2023.
  19. 19.Hanqing Li and Diego Klabjan. Reverse prompt engineering, 2025.
  20. 20.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. A diversity-promoting objective function for neural conversation models, 2016.
  21. 21.Raymond Li, Loubna Ben Allal, Yijia Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. Taco: Topics in algorithmic code generation. arXiv preprint arXiv:2312.14852, 2023.
  22. 22.Licheng Liu, Zihan Wang, Linjie Li, Chenwei Xu, Yiping Lu, Han Liu, Avirup Sil, and Manling Li. Unary feedback as observation: Incentivizing self-reflection in large language models via multi-turn RL, 2026.
  23. 23.Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin. Understanding r1-zero-like training: A critical perspective, 2025.
  24. 24.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. Self-refine: Iterative refinement with self-feedback, 2023.
  25. 25.Justus Mattern, Sami Jaghouar, Manveer Basra, Jannik Straube, Matthew Di Ferrante, Felix Gabriel, Jack Min Ong, Vincent Weisser, and Johannes Hagemann. Synthetic-1: Two million collaboratively generated reasoning traces from deepseek-r1. https://www.primeintellect.ai/blog/synthetic-1-release, 2025. Prime Intellect dataset release.
  26. 26.Meta Llama. Llama 3.2 3b model card, 2024. Accessed 2026-01-28.
  27. 27.Ehsan Montahaei, Danial Alihosseini, and Mahdieh Soleymani Baghshah. Jointly measuring diversity and quality in text generation models, 2019.
  28. 28.John X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov, and Alexander M. Rush. Language model inversion, 2023.
  29. 29.Ted Moskovitz, Aaditya K. Singh, DJ Strouse, Tuomas Sandholm, Ruslan Salakhutdinov, Anca D. Dragan, and Stephen McAleer. Confronting reward model overoptimization with constrained rlhf, 2023.
  30. 30.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. Webgpt: Browser-assisted question-answering with human feedback, 2022.
  31. 31.Laura O’Mahony, Leo Grinsztajn, Hailey Schoelkopf, and Stella Biderman. Attributing mode collapse in the fine-tuning of large language models. In ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2024.
  32. 32.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback, 2022.
  33. 33.Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. Mauve: Measuring the gap between neural text and human text using divergence frontiers, 2021.
  34. 34.Penghui Qi, Zichen Liu, Xiangxin Zhou, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin. Defeating the training-inference mismatch via fp16, 2025.
  35. 35.Qwen Team. Qwen2.5 technical report, 2024.
  36. 36.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model, 2024.
  37. 37.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. Beyond accuracy: Behavioral testing of nlp models with checklist, 2020.
  38. 38.Joshua Romoff, Peter Henderson, Alexandre Piché, Vincent Francois-Lavet, and Joelle Pineau. Reward estimation for variance reduction in deep reinforcement learning, 2018.
  39. 39.Max-Philipp B. Schrader. gym-sokoban, 2018. Accessed 2026-01-29.
  40. 40.John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel. Trust region policy optimization, 2017.
  41. 41.John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. High-dimensional continuous control using generalized advantage estimation, 2018.
  42. 42.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms, 2017.
  43. 43.Stanislau Semeniuta, Aliaksei Severyn, and Sylvain Gelly. On accurate evaluation of gans for language generation, 2019.
  44. 44.Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024.
  45. 45.Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. Hybridflow: A flexible and efficient rlhf framework, 2024.
  46. 46.Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning, 2023.
  47. 47.Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. Ai models collapse when trained on recursively generated data. Nature, 631:755–759, 2024.
  48. 48.Noah Y. Siegel, Oana-Maria Camburu, Nicolas Heess, and Maria Perez-Ortiz. The probabilities also matter: A more faithful metric for faithfulness of free-text explanations in large language models, 2024.
  49. 49.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. Learning to summarize from human feedback, 2022.
  50. 50.Hanshi Sun, Momin Haider, Ruiqi Zhang, Huitao Yang, Jiahao Qiu, Ming Yin, Mengdi Wang, Peter Bartlett, and Andrea Zanette. Fast best-of-n decoding via speculative rejection, 2024.
  51. 51.Sijun Tan, Michael Luo, Colin Cai, Tarun Venkat, Kyle Montgomery, Aaron Hao, Tianhao Wu, Arnav Balyan, Manan Roongta, Chenguang Wang, Li Erran Li, Raluca Ada Popa, and Ion Stoica. rllm: A framework for post-training language agents. https://pretty-radio-b75.notion.site/rLLM-A-Framework-for-Post-Training-Language-Agents-21b81902c146819db63cd98a54ba5f31, 2025. Notion Blog.
  52. 52.Leitian Tao, Ilia Kulikov, Swarnadeep Saha, Tianlu Wang, Jing Xu, Sharon Li, Jason E Weston, and Ping Yu. Hybrid reinforcement: When reward is sparse, it’s better to be dense, 2025.
  53. 53.Guy Tevet and Jonathan Berant. Evaluating the evaluation of diversity in natural language generation, 2021.
  54. 54.Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting, 2023.
  55. 55.Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. Solving math word problems with process- and outcome-based feedback, 2022.
  56. 56.Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models, 2023.
  57. 57.Jiawei Wang, Jiacai Liu, Yuqian Fu, Yingru Li, Xintao Wang, Yuan Lin, Yu Yue, Lin Zhang, Yang Wang, and Ke Wang. Harnessing uncertainty: Entropy-modulated policy gradients for long-horizon llm agents, 2025.
  58. 58.Kangrui Wang, Pingyue Zhang, Zihan Wang, Yaning Gao, Linjie Li, Qineng Wang, Hanyang Chen, Chi Wan, Yiping Lu, Zhengyuan Yang, Lijuan Wang, Ranjay Krishna, Jiajun Wu, Fei-Fei Li, Yejin Choi, and Manling Li. VAGEN: Reinforcing world model reasoning for multi-turn VLM agents. arXiv preprint arXiv:2510.16907, 2025.
  59. 59.Ruiyi Wang and Prithviraj Ammanabrolu. A practitioner’s guide to multi-turn agentic reinforcement learning, 2025.
  60. 60.Zihan Wang, Kangrui Wang, Qineng Wang, Pingyue Zhang, Linjie Li, Zhengyuan Yang, Xing Jin, Kefan Yu, Minh Nhat Nguyen, Licheng Liu, Eli Gottlieb, Yiping Lu, Kyunghyun Cho, Jiajun Wu, Li Fei-Fei, Lijuan Wang, Yejin Choi, and Manling Li. Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning, 2025.
  61. 61.Tong Wei, Yijun Yang, Junliang Xing, Yuanchun Shi, Zongqing Lu, and Deheng Ye. Gtr: Guided thought reinforcement prevents thought collapse in rl-based vlm agent training, 2025.
  62. 62.Wujiang Xu, Wentian Zhao, Zhenting Wang, Yu-Jhe Li, Can Jin, Mingyu Jin, Kai Mei, Kun Wan, and Dimitris N. Metaxas. Epo: Entropy-regularized policy optimization for llm agents reinforcement learning, 2025.
  63. 63.Jian Yao, Ran Cheng, Xingyu Wu, Jibin Wu, and Kay Chen Tan. Diversity-aware policy optimization for large language model reasoning, 2025.
  64. 64.Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents. In Advances in Neural Information Processing Systems, volume 35, pages 20744–20757, 2022.
  65. 65.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models, 2023.
  66. 66.Baiqiao Yin, Qineng Wang, Pingyue Zhang, Jianshu Zhang, Kangrui Wang, Zihan Wang, Jieyu Zhang, Keshigeyan Chandrasegaran, Han Liu, Ranjay Krishna, Saining Xie, Manling Li, Jiajun Wu, and Li Fei-Fei. Spatial mental modeling from limited views. arXiv preprint arXiv:2506.21458, 2025.
  67. 67.Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. Metamath: Bootstrap your own mathematical questions for large language models, 2023.
  68. 68.Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, Xin Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang, Mofan Zhang, Wang Zhang, Hang Zhu, Jinhua Zhu, Jiaze Chen, Jiangjie Chen, Chengyi Wang, Hongli Yu, Yuxuan Song, Xiangpeng Wei, Hao Zhou, Jingjing Liu, Wei-Ying Ma, Ya-Qin Zhang, Lin Yan, Mu Qiao, Yonghui Wu, and Mingxuan Wang. Dapo: An open-source llm reinforcement learning system at scale, 2025.
  69. 69.Longfei Yun, Chenyang An, Zilong Wang, Letian Peng, and Jingbo Shang. The price of format: Diversity collapse in llms, 2025.
  70. 70.Kerem Zaman and Shashank Srivastava. Is chain-of-thought really not explainability? chain-of-thought can be faithful without hint verbalization, 2025.
  71. 71.Collin Zhang, John X. Morris, and Vitaly Shmatikov. Extracting prompts by inverting llm outputs, 2024.
  72. 72.Yaxiang Zhang, Yingru Li, Jiacai Liu, Jiawei Xu, Ziniu Li, Qian Liu, and Haoyuan Li. Beyond precision: Training-inference mismatch is an optimization problem and simple lr scheduling fixes it, 2026.
  73. 73.Kaijie Zhu, Qinlin Zhao, Hao Chen, Jindong Wang, and Xing Xie. Promptbench: A unified library for evaluation of large language models, 2024.
  74. 74.Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. Texygen: A benchmarking platform for text generation models, 2018.

Citation

MLA
Wang, Z., et al. “RAGEN-2: Reasoning Collapse in Agentic RL”. arXiv, 2026, https://doi.org/10.48550/arxiv.2604.06268.
APA
Wang, Z., Gui, C., Jin, X., Wang, Q., Liu, L., Wang, K., Chen, S., Li, L., Yang, Z., Zhang, P., Lu, Y., Wu, J., Fei-Fei, L., Wang, L., Choi, Y., & Li, M. (2026). RAGEN-2: Reasoning Collapse in Agentic RL. arXiv. https://doi.org/10.48550/arxiv.2604.06268
Chicago
Wang, Z., C. Gui, X. Jin, et al. 2026. “RAGEN-2: Reasoning Collapse in Agentic RL”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2604.06268.
Harvard
Wang, Z. et al. (2026) “RAGEN-2: Reasoning Collapse in Agentic RL”. arXiv. Available at: https://doi.org/10.48550/arxiv.2604.06268.
Vancouver
1. Wang Z, Gui C, Jin X, et al (2026) RAGEN-2: Reasoning Collapse in Agentic RL. https://doi.org/10.48550/arxiv.2604.06268

BibTeX

@misc{https://doi.org/10.48550/arxiv.2604.06268,
  doi = {10.48550/ARXIV.2604.06268},
  url = {https://arxiv.org/abs/2604.06268},
  author = {Wang, Zihan and Gui, Chi and Jin, Xing and Wang, Qineng and Liu, Licheng and Wang, Kangrui and Chen, Shiqi and Li, Linjie and Yang, Zhengyuan and Zhang, Pingyue and Lu, Yiping and Wu, Jiajun and Fei-Fei, Li and Wang, Lijuan and Choi, Yejin and Li, Manling},
  keywords = {Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {RAGEN-2: Reasoning Collapse in Agentic RL},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/