Built independently by an author, for readers. Read the story and support ChapterPal

keyword

training-based search agents

Training-based search agents are artificial intelligence systems, typically powered by large language models, whose information retrieval, browsing, and synthesis behaviors are directly optimized through model training methods such as reinforcement learning or fine-tuning rather than relying solely on prompt engineering or fixed heuristic rules. By learning from interactive feedback within web environments or document corpora, these agents learn how to autonomously formulate search queries, navigate complex web pages, filter out irrelevant or noisy content, and aggregate multi-source findings. This end-to-end optimization enables the agents to develop advanced reasoning strategies, including multi-step planning, iterative query refinement, self-correction, and source cross-validation to resolve complex, open-ended research tasks.

1 item

DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Yuxiang Zheng, Dayuan Fu, Xiangkun Hu, Xiaojie Cai, Lyumanshan Ye, Pengrui Lu, Pengfei Liu

OrganizationsGenerative AI Research Lab (GAIR)Shanghai Innovation InstituteShanghai Jiao Tong University

Why you should read this

Presents DeepResearcher, an end-to-end reinforcement learning framework that trains language model agents directly on live web search interactions to outperform prompt-based and retrieval-augmented methods on open-domain research tasks.

Large Language Models (LLMs) with web search capabilities show significant potential for deep research, yet current methods—brittle prompt engineering or RAG-based reinforcement learning in controlled environments—fail to capture real-world complexities. In this paper, we introduce DeepResearcher, the first comprehensive framework for end-to-end training of LLM-based deep research agents through scaling reinforcement learning (RL) in real-world environments with authentic web search interactions. Unlike RAG approaches reliant on fixed corpora, DeepResearcher trains agents to navigate the noisy, dynamic open web. We implement a specialized multi-agent architecture where browsing agents extract relevant information from various webpage structures and overcoming significant technical challenges. Extensive experiments on open-domain research tasks demonstrate that DeepResearcher achieves substantial improvements of up to 28.9 points over prompt engineering-based baselines and up to 7.2 points over RAG-based RL agents. Our qualitative analysis reveals emergent cognitive behaviors from end-to-end RL training, such as planning, cross-validation, self-reflection for research redirection, and maintain honesty when unable to find definitive answers. Our results highlight that end-to-end training in real-world web environments is fundamental for developing robust research capabilities aligned with real-world applications. The source code for DeepResearcher is released at: https://github.com/GAIR-NLP/DeepResearcher.

Added

2026-09-28