DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Yuxiang ZhengDayuan FuXiangkun HuXiaojie CaiLyumanshan YePengrui LuPengfei Liu

article2025EMNLP264 citations

Presents DeepResearcher, an end-to-end reinforcement learning framework that trains language model agents directly on live web search interactions to outperform prompt-based and retrieval-augmented methods on open-domain research tasks.

Listen

Modern artificial intelligence systems increasingly rely on external search tools to conduct complex research, yet current development methods suffer from severe real-world limitations. Many existing approaches rely on rigid, human-engineered prompt workflows or train models using reinforcement learning on static, simulated text databases. These controlled environments fail to prepare models for the noise, rate limits, anti-crawling defenses, and structural unpredictability of the live web. The article addresses this gap by introducing DeepResearcher, a framework designed for the end-to-end training of language model agents directly within dynamic, real-world web environments using scaled reinforcement learning.

The research evaluates whether language models trained directly against live search engines and web pages can autonomously acquire robust deep-research strategies. To achieve this, the authors curated a clean training set of 80,000 open-domain question-answering examples, filtering out subjective queries and instances where the base model already knew the answer. The system was trained end-to-end using an outcome-driven reinforcement learning algorithm on a 7-billion-parameter language model. To overcome operational bottlenecks during training, the authors deployed a 50-node server cluster to handle high-concurrency web requests, implemented API caching and retry mechanisms, and integrated specialized browsing agents that parse and extract information from web page segments in parallel.

The findings show substantial performance improvements across multiple benchmarks. DeepResearcher achieved accuracy gains of up to 28.9 points over traditional prompt-engineered baselines and outperformed previous reinforcement-learning search agents by up to 7.2 points on out-of-domain benchmarks. In an ablation comparison, training the exact same model architecture on a static local text repository resulted in dramatic performance drops, demonstrating that exposure to live web dynamics is essential for generalizability. Furthermore, the model autonomously developed critical cognitive behaviors without explicit supervision, including multi-step planning, verifying answers across independent sources, redirecting failed search strategies, and declining to answer when definitive information could not be verified.

These results carry important practical implications for organizations developing autonomous intelligence tools. Direct end-to-end training in authentic operational environments eliminates the need for fragile, human-crafted heuristic pipelines while significantly enhancing reliability and generalization. The emergence of cross-validation and self-restraint behaviors also reduces the risk of automated hallucinations in high-stakes research workflows. For decision-makers, the article demonstrates that investing in high-concurrency infrastructure to train agents in live operating conditions delivers superior, more robust autonomous reasoning capabilities than relying on curated, synthetic testbeds.

Moving forward, the authors recommend expanding this framework to substantially larger base models to determine if further reasoning scaling can be achieved. Organizations should also consider developing more nuanced reward mechanisms tailored to unstructured, long-form synthesis rather than short, factual answers. While the reported results provide high confidence in the viability of live reinforcement learning for search agents, caution is advised regarding operational safeguards, as deploying autonomous browsing at scale requires robust rate limiting and strict compliance with ethical and privacy standards.

arXiv: 2504.03160GAIR-NLP/DeepResearcher
Cover for DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Abstract

Large Language Models (LLMs) with web search capabilities show significant potential for deep research, yet current methods—brittle prompt engineering or RAG-based reinforcement learning in controlled environments—fail to capture real-world complexities. In this paper, we introduce DeepResearcher, the first comprehensive framework for end-to-end training of LLM-based deep research agents through scaling reinforcement learning (RL) in real-world environments with authentic web search interactions. Unlike RAG approaches reliant on fixed corpora, DeepResearcher trains agents to navigate the noisy, dynamic open web. We implement a specialized multi-agent architecture where browsing agents extract relevant information from various webpage structures and overcoming significant technical challenges. Extensive experiments on open-domain research tasks demonstrate that DeepResearcher achieves substantial improvements of up to 28.9 points over prompt engineering-based baselines and up to 7.2 points over RAG-based RL agents. Our qualitative analysis reveals emergent cognitive behaviors from end-to-end RL training, such as planning, cross-validation, self-reflection for research redirection, and maintain honesty when unable to find definitive answers. Our results highlight that end-to-end training in real-world web environments is fundamental for developing robust research capabilities aligned with real-world applications. The source code for DeepResearcher is released at: https://github.com/GAIR-NLP/DeepResearcher.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Prompt-Based Search Agents
  • 2.2 Training-Based Search Agents
  • 2.3 Training Environments
  • 3 Methodology
  • 3.1 Deep Research Trajectory
  • 3.2 Addressing Challenges in Dynamic Real-World Web Environments
  • 3.3 RL Training Framework
  • 3.4 Reward
  • 4 Experiments
  • 4.1 Experimental Setups
  • 4.1.1 Training Data Curation
  • 4.1.2 Model and Hyperparameters
  • 4.2 Evaluation and Results
  • 4.2.1 Benchmarks
  • 4.2.2 Baselines
  • 4.2.3 Evaluation Metrics
  • 4.2.4 Main Results
  • 5 Analysis
  • 5.1 Training Dynamics
  • 5.2 Case Study
  • 6 Conclusion
  • Limitations
  • Ethical Considerations
  • Acknowledgments
  • References
  • A Beyond Memorization: Curating Search-Dependent Training Data
  • A.1 Leveraging Open Domain QA Data
  • A.2 The Issue of Data Contamination
  • A.3 Data Cleaning and Contamination Detection
  • B Case Study Example
  • C Prompts
  • C.1 Prompt for Question Quality Level Evaluation
  • C.2 Prompt for Model's Answer Quality Level Evaluation
  • C.3 Prompt for Research Plan on Question Answering
  • D Training Scaling Result
  • E Performance

Knowls

  1. Knowl 1 — DeepResearcher Agent Architecture and Interaction Pipeline

    model/method

    DeepResearcher is an autonomous deep research agent trained to interact directly with real-world, dynamic web environments using iterative reasoning and tool invocation.

    The trajectory of the primary research agent proceeds in structured cycles until sufficient information is acquired:

    1. Reasoning: Prior to any action, the model performs chain-of-thought deliberation enclosed within <think>...</think> tags.
    2. Tool Selection and Invocation: When external knowledge is required, the agent issues a JSON-formatted request containing tool names and parameters inside <tool_call>...</tool_call> tags.
      • web_search: Issues search queries to a search engine API (e.g., Google Search), returning structured results comprising titles, URLs, and text snippets for the top-kk retrieved pages (default k=10k=10).
      • browse_webpage: Invokes a multi-agent browsing subsystem to inspect specific URLs. The browsing tool partitions lengthy webpages into segments and deploys parallel Reading Agents that process pages sequentially from the first segment, maintaining a short-term memory buffer and deciding whether to read subsequent segments or terminate browsing if the content is irrelevant. A Synthesis Agent then aggregates and summarizes findings across all processed URLs into a unified observation returned to the primary agent.
    3. Observation Processing: Tool responses are fed back into the agent context.
    4. Final Answer: Once the agent deems the accumulated evidence sufficient, it produces the final answer enclosed within <answer>...</answer> tags.
  2. Knowl 2 — Group Relative Policy Optimization with Observation Masking for Deep Research

    model/method

    DeepResearcher trains an LLM agent end-to-end using Group Relative Policy Optimization (GRPO) to discover search, browsing, and reasoning strategies without relying on supervised demonstration priors or human-crafted workflows.

    For an input question x∼Dx \sim \mathcal{D} drawn from the training distribution D\mathcal{D}, the policy generates a group of GG rollout trajectories: τ={yi}i=1G∼πθold(⋅∣x)\tau = \{y_i\}_{i=1}^G \sim \pi_{\theta_{\text{old}}}(\cdot|x)

    Instead of maintaining a separate critic network, GRPO calculates the advantage AiA_i for each rollout trajectory yiy_i relative to the mean and standard deviation of rewards within the group of GG trajectories. The policy πθ\pi_\theta is updated by maximizing the clipped surrogate objective with a Kullback-Leibler (KL) divergence penalty against a reference policy πθref\pi_{\theta_{\text{ref}}}: J(θ)=Ex∼D,{yi}i=1G∼πθold(⋅∣x)[1G∑i=1G(min⁡(πθ(yi∣x)πθold(yi∣x)Ai,clip(πθ(yi∣x)πθold(yi∣x),1−ϵ,1+ϵ)Ai)−βDKL(πθ∥πθref))]J(\theta) = \mathbb{E}_{x \sim \mathcal{D}, \{y_i\}_{i=1}^G \sim \pi_{\theta_{\text{old}}}(\cdot|x)} \left[ \frac{1}{G} \sum_{i=1}^G \left( \min\left( \frac{\pi_\theta(y_i|x)}{\pi_{\theta_{\text{old}}}(y_i|x)} A_i, \text{clip}\left(\frac{\pi_\theta(y_i|x)}{\pi_{\theta_{\text{old}}}(y_i|x)}, 1 - \epsilon, 1 + \epsilon\right) A_i \right) - \beta D_{\text{KL}}\left(\pi_\theta \parallel \pi_{\theta_{\text{ref}}}\right) \right) \right] where ϵ\epsilon is the clipping parameter and β\beta controls the KL penalty strength.

    Observation Masking: Because external tool outputs (search engine snippets, crawled webpage contents) represent environmental observations rather than tokens the policy is expected to generate, token loss masking is applied to all tool observation tokens during policy gradient computation. Loss and gradients are computed exclusively on tokens generated directly by the agent (reasoning tokens, tool call specifications, and final answer tokens).

  3. Knowl 3 — Outcome-Based Reward Formulation for DeepResearcher

    equation

    DeepResearcher is optimized using purely outcome-based rewards computed at the end of each trajectory, penalizing structural format violations and evaluating answer quality via token-level overlap against ground-truth answers: reward={−1if format is incorrectF1 scoreif format is correct\text{reward} = \begin{cases} -1 & \text{if format is incorrect} \\ \text{F1 score} & \text{if format is correct} \end{cases}

    where:

    • Format Penalty: Assigns a fixed scalar reward of −1-1 if the output trajectory violates expected structural formats (e.g., missing <think>, <tool_call>, or <answer> tags, or generating invalid JSON syntax in tool parameters).
    • F1 Reward: For syntactically valid outputs, the reward equals the word-level F1 score (harmonic mean of precision and recall over normalized tokens) between the text extracted from the final <answer> tag and the ground-truth short answer. Both strings are lowercased and stripped of punctuation before scoring.
  4. Knowl 4 — Distributed Infrastructure and API Management for Real-Web RL Scaling

    model/method

    Scaling reinforcement learning with real-world web search environments introduces critical engineering bottlenecks: massive I/O concurrency during parallel rollouts, search provider rate limits (e.g., 200 queries/second), anti-crawling countermeasures, and network latency. DeepResearcher resolves these through three key mechanisms:

    1. Distributed I/O Server Cluster: A 50-node CPU cluster specifically dedicated to handling tool execution requests generated across 4,096 parallel rollout trajectories per training step. The cluster dispatches search API requests, manages parallel HTTP crawling pipelines, and processes extracted DOM text.
    2. Exception and Anti-Crawl Retry Pipeline: Automated retry mechanisms with exponential backoff handle blocked requests, non-responsive web servers, and HTTP status errors encountered during page crawling or API access.
    3. 7-Day Query-Result Caching: A centralized caching system stores the returned structured responses for every search query. If an identical search query is issued within a 7-day window, results are served directly from cache, mitigating API rate-limit bottlenecks and substantially reducing commercial API query costs.
  5. Knowl 5 — Two-Stage Search-Dependent Training Data Curation and Decontamination

    model/method

    To prevent models from relying on parametric memory and ensure they learn active search behaviors, training datasets undergo a two-stage curation and contamination detection pipeline:

    1. Low-Quality and Undesirable Question Filtering: DeepSeek-R1 is prompted to evaluate candidate questions and filter out:
      • Time-sensitive questions whose ground truth shifts over time (e.g., "Who is the current CEO of Apple?").
      • Subjective questions lacking factual consensus (e.g., "What is the best smartphone?").
      • Harmful or policy-violating queries.
    2. Contamination Detection via Pass@10 Screening: To filter out questions already stored in the base model's parametric knowledge, 10 independent responses are sampled from the un fine-tuned base LLM (Qwen2.5-7B-Instruct) without web access. If any of the 10 generations contains the ground-truth answer (i.e., pass@10>0\text{pass}@10 > 0), the question is classified as contaminated and removed from the training set.

    Applying this pipeline across open-domain QA sources yields a 80,000-sample training dataset with a 1:1:3:3 ratio across NaturalQuestions (NQ), TriviaQA (TQ), HotpotQA, and 2WikiMultiHopQA (2Wiki), deliberately allocating 75% of the data to multi-hop reasoning tasks.

  6. Knowl 6 — In-Domain Performance of DeepResearcher Across Open-Domain QA Benchmarks

    data/table

    The in-domain evaluation compares DeepResearcher (using Qwen2.5-7B-Instruct trained with GRPO and real web search) against prompt-based baselines, local RAG RL baselines, and web search baselines across 512 randomly sampled development instances from NaturalQuestions (NQ), TriviaQA (TQ), HotpotQA, and 2WikiMultiHopQA (2Wiki). Performance is measured by word-level F1 and Model-Based Evaluation (MBE) accuracy judged by GPT-4o-mini.

    Method Inference Env. NQ TQ HotpotQA 2Wiki
    F1 MBE F1 MBE F1 MBE F1 MBE
    Prompt Based
    CoT Local RAG 19.8 32.0 45.6 48.2 24.4 27.9 26.4 27.3
    CoT + RAG Local RAG 42.0 59.6 68.9 75.8 37.1 43.8 24.4 24.8
    Search-o1* Local RAG 34.5 57.4 52.6 61.1 31.6 40.8 28.6 32.8
    Search-o1 Web Search 32.4 55.1 58.9 69.5 33.0 42.4 30.9 37.7
    ReAct-style Agent Web Search 22.7 39.6 41.9 49.2 19.7 26.2 17.6 17.6
    Training Based
    Search-r1-base Local RAG 45.4 60.0 71.9 76.2 55.9 63.0 44.6 47.9
    Search-r1-instruct Local RAG 33.1 49.6 44.7 49.2 45.7 52.5 43.4 48.8
    R1-Searcher Web Search 35.4 52.3 73.1 79.1 44.8 53.1 59.4 65.8
    DeepResearcher (Local RAG) Local RAG 29.5 46.3 51.9 55.5 29.4 35.4 26.3 27.5
    DeepResearcher Web Search 39.6 61.9 78.4 85.0 52.8 64.3 59.7 66.6

    DeepResearcher achieves the highest MBE across all four datasets (61.9 on NQ, 85.0 on TQ, 64.3 on HotpotQA, and 66.6 on 2Wiki). The ablation agent trained strictly on a static local RAG repository ("DeepResearcher (Local RAG)") suffers large performance drops across every metric, demonstrating that live web search exposure during training is essential for robust research policies.

  7. Knowl 7 — Out-of-Domain Generalization Performance of DeepResearcher

    data/table

    Out-of-domain (OOD) generalization was evaluated on MuSiQue (512 dev samples), Bamboogle (all 125 dev samples, containing questions not resolvable strictly via Wikipedia), and PopQA (512 dev samples). Methods were evaluated using normalized word-level F1 and GPT-4o-mini Model-Based Evaluation (MBE) accuracy.

    Method Inference Env. MuSiQue Bamboogle PopQA
    F1 MBE F1 MBE F1 MBE
    Prompt Based
    CoT Local RAG 8.5 7.4 22.1 21.6 17.0 15.0
    CoT + RAG Local RAG 10.0 10.0 25.4 27.2 46.9 48.8
    Search-o1* Local RAG 16.8 21.3 35.8 38.4 36.9 42.4
    Search-o1 Web Search 14.7 19.7 46.6 53.6 38.3 43.4
    ReAct-style Agent Web Search 8.9 10.0 34.4 36.8 19.1 20.5
    Training Based
    Search-r1-base Local RAG 26.7 27.5 56.5 57.6 43.2 47.0
    Search-r1-instruct Local RAG 26.5 28.3 45.0 47.2 43.0 44.5
    R1-Searcher Web Search 22.8 25.6 64.8 65.6 42.7 43.4
    DeepResearcher (Local RAG) Local RAG 12.7 12.5 42.7 46.4 23.2 23.4
    DeepResearcher Web Search 27.1 29.3 71.0 72.8 48.5 52.7

    DeepResearcher outperforms all baselines across all three out-of-domain benchmarks. On Bamboogle, DeepResearcher attains 71.0 F1 / 72.8 MBE, outperforming R1-Searcher (64.8 F1 / 65.6 MBE) and local RAG models by a large margin.

  8. Knowl 8 — Emergent Cognitive Behaviors from End-to-End Real-Web RL

    empirical result

    Qualitative inspection of DeepResearcher's trajectories throughout GRPO training reveals four emergent cognitive behaviors that arise without explicit supervised fine-tuning (SFT) demonstrations:

    1. Autonomous Planning: For multi-hop questions, the agent breaks the problem down into sequential steps in <think> (e.g., Step 1: Identify entity, Step 2: Find entity's location, Step 3: Find landmark), dynamically merging or modifying sub-goals as new information arrives.
    2. Cross-Validation: Even after retrieving candidate answers from an initial search, the agent executes follow-up queries or page visits across independent sources to corroborate findings before committing to a final response.
    3. Reflection and Directional Adjustment: When search observations return irrelevant, noisy, or unexpected information (e.g., name ambiguities), the agent recognizes the discrepancy during reasoning and reformulates query keywords in subsequent calls.
    4. Honesty and Limitation Acknowledgment: When extensive web search and page browsing fail to locate definitive, specific facts (e.g., exact city-level statistics when only country-level data exists), the agent states its inability to produce a precise number rather than hallucinating.
  9. Knowl 9 — Training Dynamics of RL Scaling in Deep Research

    empirical result

    Analysis of DeepResearcher's training progression under GRPO demonstrates consistent scaling characteristics:

    • Continuous Performance Growth: Overall average F1 increases smoothly from approximately 0.375 at step 0 to ~0.55 after 30+ training steps across benchmarks.
    • Adaptive Tool-Call Scaling: The average number of tool calls per trajectory increases with problem difficulty. While simpler 1-hop and 2-hop questions plateau early, 4-hop questions exhibit an uninterrupted upward trajectory in tool invocations beyond 34 steps.
    • Response Length Expansion: Average trajectory response lengths consistently expand as training progresses across all reasoning hop levels (reaching 5,000–5,500 tokens for 4-hop questions) as the model generates more extensive planning, verification, and reflection tokens without hitting saturation.
  10. Knowl 10 — Limitations of DeepResearcher

    limitation

    The DeepResearcher framework has two main limitations:

    1. Base Model Scale Restriction: Experiments were conducted exclusively with a 7-billion parameter backbone (Qwen2.5-7B-Instruct). The scaling behavior, emergent research capabilities, and performance gains on significantly larger parameter regimes (e.g., 70B+ LLMs) remain unexplored.
    2. Reward Metric Limitation for Open-Ended Research: The training relies on word-level F1 matching against short, factual ground-truth answers paired with structural format penalties. This outcome-based reward is unsuited for open-ended, ill-defined deep research inquiries that require generating multi-page, synthesized long-form reports where exact string overlap metrics do not apply.

Coverage note — None was omitted; all contributed methods, training objectives, infrastructure implementations, data curation protocols, empirical evaluation results, and stated limitations are covered.

References

  1. 1.Salaheddin Alzubi, Creston Brooks, Purva Chiniya, Edoardo Contente, Chiara von Gerlach, Lucas Irwin, Yihan Jiang, Arda Kaz, Windsor Nguyen, Sewoong Oh, and 1 others. 2025. Open deep search: Democratizing search with open-source reasoning agents. arXiv preprint arXiv:2503.20201.
  2. 2.CAMEL-AI.org. 2025. Owl: Optimized workforce learning for general multi-agent assistance in real-world task automation. https://github.com/camel-ai/owl. Accessed: 2025-03-07.
  3. 3.Mingyang Chen, Tianpeng Li, Haoze Sun, Yijie Zhou, Chenzheng Zhu, Haofen Wang, Jeff Z. Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen. 2025. Research: Learning to reason with search for llms via reinforcement learning. Preprint, arXiv:2503.19470.
  4. 4.Wenfeng Feng, Chuzhan Hao, Yuewei Zhang, Jingyi Song, and Hao Wang. 2025. Airrag: Activating intrinsic reasoning for retrieval augmented generation via tree-based search. arXiv preprint arXiv:2501.10053.
  5. 5.Dayuan Fu, Keqing He, Yejie Wang, Wentao Hong, Zhuoma Gongque, Weihao Zeng, Wei Wang, Jingang Wang, Xunliang Cai, and Weiran Xu. 2025. Agentrefine: Enhancing agent generalization through refinement tuning. arXiv preprint arXiv:2501.01702.
  6. 6.Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997, 2.
  7. 7.Google. 2024. Gemini deep research.
  8. 8.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, and 1 others. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948.
  9. 9.Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps. In Proceedings of the 28th International Conference on Computational Linguistics, pages 6609–6625, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  10. 10.Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. 2024. MetaGPT: Meta programming for a multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations.
  11. 11.Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and 1 others. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276.
  12. 12.Bowen Jin, Hansi Zeng, Zhenrui Yue, Dong Wang, Hamed Zamani, and Jiawei Han. 2025. Search-r1: Training llms to reason and leverage search engines with reinforcement learning. arXiv preprint arXiv:2503.09516.
  13. 13.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  14. 14.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  15. 15.Xiaoxi Li, Guanting Dong, Jiajie Jin, Yuyao Zhang, Yujia Zhou, Yutao Zhu, Peitian Zhang, and Zhicheng Dou. 2025a. Search-o1: Agentic search-enhanced large reasoning models. arXiv preprint arXiv:2501.05366.
  16. 16.Xuefeng Li, Haoyang Zou, and Pengfei Liu. 2025b. Limr: Less is more for rl scaling. Preprint, arXiv:2502.11886.
  17. 17.Xuefeng Li, Haoyang Zou, and Pengfei Liu. 2025c. Torl: Scaling tool-integrated rl. Preprint, arXiv:2503.23383.
  18. 18.Xinbin Liang, Jinyu Xiang, Zhaoyang Yu, Jiayi Zhang, and Sirui Hong. 2025. Openmanus: An open-source framework for building general ai agents. https://github.com/mannaandpoem/OpenManus.
  19. 19.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, and Daniel Khashabi. 2022. When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories. arXiv preprint.
  20. 20.OpenAI. 2024. Learning to reason with llms, september 2024.
  21. 21.OpenAI. 2025. Deep research system card. Technical report, OpenAI.
  22. 22.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. 2022. Measuring and narrowing the compositionality gap in language models. arXiv preprint arXiv:2210.03350.
  23. 23.Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, and 1 others. 2023. Toolllm: Facilitating large language models to master 16000+ real-world apis. arXiv preprint arXiv:2307.16789.
  24. 24.Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others. 2025. Qwen2.5 technical report. Preprint, arXiv:2412.15115.
  25. 25.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36:68539–68551.
  26. 26.Huatong Song, Jinhao Jiang, Yingqian Min, Jie Chen, Zhipeng Chen, Wayne Xin Zhao, Lei Fang, and Ji-Rong Wen. 2025. R1-searcher: Incentivizing the search capability in llms via reinforcement learning. arXiv preprint arXiv:2503.05592.
  27. 27.Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, and 1 others. 2025. Kimi k1.5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599.
  28. 28.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multi-hop questions via single-hop question composition. Transactions of the Association for Computational Linguistics.
  29. 29.Prakhar Verma, Sukruta Prakash Midigeshi, Gaurav Sinha, Arno Solin, Nagarajan Natarajan, and Amit Sharma. 2025. Plan*rag: Efficient test-time planning for retrieval augmented generation. Preprint, arXiv:2410.20753.
  30. 30.Xiaohua Wang, Zhenghua Wang, Xuan Gao, Feiran Zhang, Yixin Wu, Zhibo Xu, Tianyuan Shi, Zhengyuan Wang, Shizheng Li, Qi Qian, and 1 others. 2024a. Searching for best practices in retrieval-augmented generation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 17716–17736.
  31. 31.Ziting Wang, Haitao Yuan, Wei Dong, Gao Cong, and Feifei Li. 2024b. Corag: A cost-constrained retrieval optimization system for retrieval-augmented generation. arXiv preprint arXiv:2411.00744.
  32. 32.xAI. 2025. Grok 3.
  33. 33.Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2024. Alignment for honesty. Advances in Neural Information Processing Systems, 37:63565–63598.
  34. 34.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Conference on Empirical Methods in Natural Language Processing (EMNLP).
  35. 35.Tian Yu, Shaolei Zhang, and Yang Feng. 2024. Auto-rag: Autonomous retrieval-augmented generation for large language models. arXiv preprint arXiv:2411.19443.
  36. 36.Murong Yue, Wenlin Yao, Haitao Mi, Dian Yu, Ziyu Yao, and Dong Yu. 2024a. Dots: Learning to reason dynamically in llms via optimal reasoning trajectories search. arXiv preprint arXiv:2410.03864.
  37. 37.Zhenrui Yue, Honglei Zhuang, Aijun Bai, Kai Hui, Rolf Jagerman, Hansi Zeng, Zhen Qin, Dong Wang, Xuanhui Wang, and Michael Bendersky. 2024b. Inference scaling for long-context retrieval augmented generation. arXiv preprint arXiv:2410.04343.
  38. 38.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023.
  39. 39.Yuxiang Zheng, Shichao Sun, Lin Qiu, Dongyu Ru, Cheng Jiayang, Xuefeng Li, Jifan Lin, Binjie Wang, Yun Luo, Renjie Pan, Yang Xu, Qingkai Min, Zizhao Zhang, Yiwen Wang, Wenjie Li, and Pengfei Liu. 2024. OpenResearcher: Unleashing AI for accelerated scientific research. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 209–218, Miami, Florida, USA. Association for Computational Linguistics.

Citation

MLA
Zheng, Y., et al. “DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 414–31, https://doi.org/10.18653/v1/2025.emnlp-main.22.
APA
Zheng, Y., Fu, D., Hu, X., Cai, X., Ye, L., Lu, P., & Liu, P. (2025). DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 414–431. https://doi.org/10.18653/v1/2025.emnlp-main.22
Chicago
Zheng, Y., D. Fu, X. Hu, et al. 2025. “DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 414–31. https://doi.org/10.18653/v1/2025.emnlp-main.22.
Harvard
Zheng, Y. et al. (2025) “DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments”, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 414–431. Available at: https://doi.org/10.18653/v1/2025.emnlp-main.22.
Vancouver
1. Zheng Y, Fu D, Hu X, Cai X, Ye L, Lu P, Liu P (2025) DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 414–431

BibTeX

@inproceedings{zheng-etal-2025-deepresearcher,
    title = "{D}eep{R}esearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments",
    author = "Zheng, Yuxiang  and
      Fu, Dayuan  and
      Hu, Xiangkun  and
      Cai, Xiaojie  and
      Ye, Lyumanshan  and
      Lu, Pengrui  and
      Liu, Pengfei",
    editor = "Christodoulopoulos, Christos  and
      Chakraborty, Tanmoy  and
      Rose, Carolyn  and
      Peng, Violet",
    booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2025",
    address = "Suzhou, China",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.emnlp-main.22/",
    doi = "10.18653/v1/2025.emnlp-main.22",
    pages = "414--431",
    ISBN = "979-8-89176-332-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/