DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
Yuxiang ZhengDayuan FuXiangkun HuXiaojie CaiLyumanshan YePengrui LuPengfei Liu
Presents DeepResearcher, an end-to-end reinforcement learning framework that trains language model agents directly on live web search interactions to outperform prompt-based and retrieval-augmented methods on open-domain research tasks.
Modern artificial intelligence systems increasingly rely on external search tools to conduct complex research, yet current development methods suffer from severe real-world limitations. Many existing approaches rely on rigid, human-engineered prompt workflows or train models using reinforcement learning on static, simulated text databases. These controlled environments fail to prepare models for the noise, rate limits, anti-crawling defenses, and structural unpredictability of the live web. The article addresses this gap by introducing DeepResearcher, a framework designed for the end-to-end training of language model agents directly within dynamic, real-world web environments using scaled reinforcement learning.
The research evaluates whether language models trained directly against live search engines and web pages can autonomously acquire robust deep-research strategies. To achieve this, the authors curated a clean training set of 80,000 open-domain question-answering examples, filtering out subjective queries and instances where the base model already knew the answer. The system was trained end-to-end using an outcome-driven reinforcement learning algorithm on a 7-billion-parameter language model. To overcome operational bottlenecks during training, the authors deployed a 50-node server cluster to handle high-concurrency web requests, implemented API caching and retry mechanisms, and integrated specialized browsing agents that parse and extract information from web page segments in parallel.
The findings show substantial performance improvements across multiple benchmarks. DeepResearcher achieved accuracy gains of up to 28.9 points over traditional prompt-engineered baselines and outperformed previous reinforcement-learning search agents by up to 7.2 points on out-of-domain benchmarks. In an ablation comparison, training the exact same model architecture on a static local text repository resulted in dramatic performance drops, demonstrating that exposure to live web dynamics is essential for generalizability. Furthermore, the model autonomously developed critical cognitive behaviors without explicit supervision, including multi-step planning, verifying answers across independent sources, redirecting failed search strategies, and declining to answer when definitive information could not be verified.
These results carry important practical implications for organizations developing autonomous intelligence tools. Direct end-to-end training in authentic operational environments eliminates the need for fragile, human-crafted heuristic pipelines while significantly enhancing reliability and generalization. The emergence of cross-validation and self-restraint behaviors also reduces the risk of automated hallucinations in high-stakes research workflows. For decision-makers, the article demonstrates that investing in high-concurrency infrastructure to train agents in live operating conditions delivers superior, more robust autonomous reasoning capabilities than relying on curated, synthetic testbeds.
Moving forward, the authors recommend expanding this framework to substantially larger base models to determine if further reasoning scaling can be achieved. Organizations should also consider developing more nuanced reward mechanisms tailored to unstructured, long-form synthesis rather than short, factual answers. While the reported results provide high confidence in the viability of live reinforcement learning for search agents, caution is advised regarding operational safeguards, as deploying autonomous browsing at scale requires robust rate limiting and strict compliance with ethical and privacy standards.
- Paper: WebGPT: Browser-assisted question-answering with human feedback, Reiichiro Nakano et al. (2021). WebGPT provides the foundational paradigm of training LLMs to perform browser-assisted search and evidence extraction using reinforcement learning and human feedback.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). ReAct establishes the foundational framework for interleaving verbal reasoning traces and external action execution that web browsing agents rely on.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). WebArena introduces the standard realistic web environment and evaluation benchmarks for testing autonomous agent navigation and information extraction.
- Paper: DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning, A list of authors and their affiliations appears at the end of the paper (2025). DeepSeek-R1 demonstrates how large-scale rule-based reinforcement learning with verifiable rewards elicits emergent cognitive behaviors like planning and self-reflection in language models.
- Paper: AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?, Ori Yoran et al. (2024). AssistantBench establishes the realistic multi-step web browsing challenge and analyzes why standard prompting and retrieval methods fail on live websites.
- Paper: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection, Akari Asai et al. (2024). Self-RAG formalizes how models can learn to reflect, retrieve on demand, and critique retrieved content to improve factual grounding.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). This survey provides essential background on agent architectures, memory, and planning modules for LLM-based autonomous systems.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). Search-R1 builds directly on the paradigm of training reasoning models via reinforcement learning with search engines by formalizing retrieved token loss masking for multi-turn search interactions.
- Paper: Search-o1: Agentic Search-Enhanced Large Reasoning Models, Xiaoxi Li et al. (2025). Search-o1 extends search-integrated reasoning agents by triggering active search queries dynamically during long reasoning trajectories and refining documents into concise facts.
- Paper: WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration, Yao Zhang et al. (2025). WebPilot extends multi-agent web navigation architectures by combining high-level strategic planning with localized, reflection-guided tree search in complex environments.
- Paper: Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments, Hongjin Su et al. (2025). Learn-by-interact generalizes agent training by autonomously synthesizing realistic environment tasks and trajectories via backward construction without human labeling.
- Paper: The Art of Scaling Reinforcement Learning Compute for LLMs, Devvrit Khatri et al. (2026). This work analyzes the compute scaling laws and empirical design choices underlying the reinforcement learning post-training used to scale LLM reasoning.
- Paper: Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context, Keivan Alizadeh et al. (2026). SRLM addresses long-context research and browsing synthesis challenges by using self-reflective program search over extensive information spaces.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). This comprehensive survey categorizes agentic reasoning and post-training reinforcement learning methods across foundational single-agent and multi-agent interaction layers.
