RAGEN-2: Reasoning Collapse in Agentic RL
Zihan WangChi GuiXing JinQineng WangLicheng LiuKangrui WangShiqi ChenLinjie LiZhengyuan YangPingyue Zhang
Reveals that reinforcement learning agents often succumb to input-agnostic template collapse undetected by standard entropy metrics, and introduces mutual information diagnostics alongside signal-to-noise prompt filtering to restore responsive reasoning and improve performance.
Training multi-turn artificial intelligence agents using reinforcement learning is notoriously unstable. Practitioners commonly monitor process stability using entropy, which measures output diversity. However, the article reveals that entropy only tracks diversity within the same input, failing to detect whether the model's reasoning actually adapts across different tasks and prompts. Consequently, models can experience "template collapse," where they produce superficially diverse yet completely generic, input-agnostic reasoning templates. Because this failure mode remains invisible to standard entropy and reward metrics, models can silently degrade into unreliable decision-makers.
The main objective of the article is to diagnose template collapse, explain its underlying mathematical causes, and demonstrate an effective, computationally lightweight mitigation strategy. The researchers formulate an information-theoretic framework to measure true input dependence and test an adaptive optimization method to maintain high reasoning quality.
To accomplish this, the authors evaluate large language model agents across seven complementary synthetic testbeds spanning irreversible spatial planning, stochastic grid navigation, symbolic math reasoning, web search, online shopping navigation, and competitive code generation. They analyze multiple reinforcement learning algorithms—including standard policy optimization and group-relative variants—across diverse model architectures ranging from 0.5 billion to 7 billion parameters, as well as multimodal vision-language models. To track input sensitivity online without needing external evaluation models, the authors introduce a mutual information proxy computed via in-batch cross-scoring.
The article establishes several key findings. First, mutual information proxies reliably diagnose reasoning degradation early in training, demonstrating a positive correlation with final task performance (+0.39), whereas standard entropy metrics correlate negatively (−0.11 to −0.14) and point in the wrong direction. Second, the authors explain template collapse through a signal-to-noise ratio mechanism: when within-prompt reward variance is near zero, task-discriminative gradients vanish, allowing input-agnostic regularizers (such as entropy bonuses and reference-model constraints) to dominate and erase input-specific reasoning. Third, the proposed intervention—SNR-Aware Filtering, which dynamically retains only high-reward-variance prompts before computing parameter updates—consistently improves peak performance across tasks, models, and modalities. For example, in planning and navigation benchmarks, filtering increased task success rates by up to 16 to 59 percentage points while simultaneously cutting per-step gradient computation time by 26% to 41%.
These findings imply that conventional training pipelines often waste computational resources on uninformative data that actively harms reasoning capabilities. Standard stabilization techniques, such as adjusting entropy or penalty coefficients, cannot prevent template collapse because they do not address the underlying signal-to-noise imbalance. Implementing variance-based filtering directly enhances gradient signal quality, providing a practical, cost-effective way to improve agent robustness and learning efficiency without requiring extra models or supervision.
Based on these results, engineering teams should replace or supplement entropy tracking with mutual information proxies to monitor online agent training. Training pipelines should adopt adaptive, nucleus-style reward variance filtering to eliminate low-signal updates. Practitioners can run a quick diagnostic check—measuring the dispersion ratio of reward variance across a single rollout batch—to determine beforehand whether an environment has sufficient variance heterogeneity to benefit from filtering.
While the findings are robust across diverse single-agent tasks and modalities, certain limitations apply. Reward variance acts as an effective signal proxy only when environments possess sufficient variance heterogeneity; in purely noisy or uniformly sparse environments, variance filtering offers diminished benefits. Furthermore, the framework has not yet been validated in multi-agent settings, and extremely aggressive filtering parameters could potentially narrow exploration if not tuned per task. Readers should view the method as a proven stabilizer for single-agent optimization that requires task-specific calibration of data retention rates.
- Paper: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, DeepSeek-AI et al. (2025). This paper establishes the foundational Group Relative Policy Optimization (GRPO) framework and outcome-based RL paradigm for LLM reasoning that RAGEN-2 directly builds upon and analyzes.
- Paper: Understanding R1-Zero-Like Training: A Critical Perspective, Zichen Liu et al. (2025). It provides a critical mathematical examination of GRPO and base model reasoning biases, offering crucial background for understanding the algorithmic instabilities studied in RAGEN-2.
- Paper: Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?, Yang Yue et al. (2025). It explores how RLVR affects exploration and base model capabilities, directly motivating RAGEN-2's investigation into template collapse and spurious diversity.
- Paper: SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training, Tianzhe Chu et al. (2025). It investigates generalization and memorization dynamics in post-training RL, framing the core challenge of why models collapse onto brittle reasoning patterns.
- Paper: ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness, Archiki Prasad et al. (2023). It introduces information-theoretic evaluation of reasoning chain quality beyond final accuracy, anticipating RAGEN-2's information-theoretic decomposition into diversity and distinguishability.
- Paper: Evaluating Mathematical Reasoning Beyond Accuracy, Shijie Xia et al. (2025). It analyzes structural step-level reasoning quality beyond final-answer correctness, providing useful prerequisite context for diagnosing collapsed reasoning chains.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). It lays the groundwork for interleaved reasoning and acting in LLM agents across diverse interactive environments, which serve as the primary testbed in RAGEN-2.
- Paper: Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs, Zhiyuan Hu et al. (2026). It directly complements RAGEN-2 by developing uniqueness-aware reinforcement learning rewards to actively mitigate strategy-level exploration collapse during training.
- Paper: Demystifying Reinforcement Learning Post-Training of Language Models, Donovan Clay et al. (2026). It extends the investigation of post-training RL dynamics by analyzing how prompt distributions, base priors, and reward densities interact during policy optimization.
- Paper: When Can LLMs Learn to Reason with Weak Supervision?, Salman Rahman et al. (2026). It analyzes reasoning saturation dynamics and generalization under weak supervision, offering insights into reward signal degradation related to RAGEN-2's SNR findings.
- Paper: GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization, Shih-Yang Liu et al. (2026). It addresses optimization collapse in multi-reward RL by decoupling reward normalization, providing an architectural solution to reward gradient issues explored in RAGEN-2.
- Paper: Understanding Reasoning from Pretraining to Post-Training, Jingyan Shen et al. (2026). It builds on reasoning behavior analysis by quantifying how RL alters pre-trained policy distributions across varying problem difficulty levels.
- Paper: SPIRAL: Learning to Search and Aggregate, Jubayer Ibn Hamid et al. (2026). It scales beyond single-trace RL by joint sequential, parallel, and aggregative search training, preserving reasoning entropy across generation stages.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). It provides a comprehensive survey that situates agentic RL training stability and diagnostic mechanisms within the broader landscape of agentic reasoning frameworks.
