Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Bowen JinHansi ZengZhenrui YueDong WangHamed ZamaniJiawei Han

article2025arXiv1,463 citations

Proposes Search-R1, a reinforcement learning framework that trains language models to autonomously formulate multi-turn search queries during step-by-step reasoning, outperforming traditional retrieval-augmented generation baselines by up to 41% on question-answering benchmarks.

Listen

Large language models often struggle with complex, multi-step reasoning and lack access to up-to-date factual information. While standard retrieval-augmented generation and prompting strategies allow models to query external databases, they rely on fixed retrieval pipelines or extensive manual prompting, preventing models from learning how to interact adaptively with search engines. Training-based alternatives typically require expensive, large-scale human demonstrations that are difficult to scale. Reinforcement learning offers a viable alternative to instill reasoning behaviors using simple outcome feedback, yet integrating external non-differentiable search tools directly into reinforcement learning training loops presents significant stability challenges.

The article introduces SEARCH-R1, a reinforcement learning framework designed to train language models to autonomously interleave internal reasoning with multi-turn search engine interactions. It evaluates whether a model can learn optimal query generation, evidence extraction, and self-verification using only rule-based correctness rewards at the final answer step, without requiring fine-grained human demonstrations or process-based supervision.

To evaluate this framework, the authors conducted experiments across seven benchmark question-answering datasets spanning single-topic and multi-hop reasoning tasks. The approach models the search engine as part of the external environment, allowing the model to trigger search calls via structured formatting tokens during generation. A core algorithmic mechanism, retrieved token loss masking, excludes retrieved external text from gradient updates to ensure stable optimization. The authors tested the framework across multiple model sizes (ranging from 3-billion to 14-billion parameters) and compared performance against standard retrieval pipelines, chain-of-thought prompting, supervised fine-tuning, and retrieval-free reinforcement learning.

The findings show substantial performance improvements across all tested benchmarks. SEARCH-R1 achieved average relative improvements of 24% on a 7-billion parameter model and 20% on a 3-billion parameter model compared to baseline retrieval methods, with gains scaling further on 14-billion parameter models. Analysis revealed that while instruction-tuned models converge faster initially, base foundation models trained with SEARCH-R1 reach comparable final performance, showing that reinforcement learning can independently develop structured search habits. Furthermore, masking retrieved tokens during gradient updates proved critical to prevent optimization collapse, and Proximal Policy Optimization demonstrated superior stability over group-based policy algorithms over long training horizons.

These results indicate that language models do not require curated multi-turn demonstrations to master search-augmented problem-solving. A simple outcome reward tied to final correctness is sufficient to encourage sophisticated behaviors, including query refinement and iterative self-verification. By removing the dependency on costly supervision datasets, organizations can develop more accurate, self-grounding systems at lower data collection costs while reducing factual hallucinations in knowledge-intensive domains.

Organizations implementing search-augmented reasoning systems should transition from static prompt-engineered retrieval chains to integrated reinforcement learning training. Technical teams should adopt retrieved token masking to preserve training stability and calibrate retrieval breadth carefully, as retrieving excessive passages degrades performance through noise injection. Future work should explore applying this framework to multi-modal reasoning and dynamic retrieval sizing based on model uncertainty.

The conclusions are based on controlled question-answering benchmarks using a static Wikipedia knowledge corpus and exact string matching evaluation. Confidence in the empirical gains within structured information retrieval tasks is high, but practitioners should exercise caution when deploying the framework in open-ended or highly dynamic search environments where outcome verification cannot be easily automated.

Cover for Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Abstract

Efficiently acquiring external knowledge and up-to-date information is essential for effective reasoning and text generation in large language models (LLMs). Prompting advanced LLMs with reasoning capabilities to use search engines during inference is often suboptimal, as the LLM might not fully possess the capability on how to interact optimally with the search engine. This paper introduces Search-R1, an extension of reinforcement learning (RL) for reasoning frameworks where the LLM learns to autonomously generate (multiple) search queries during step-by-step reasoning with real-time retrieval. Search-R1 optimizes LLM reasoning trajectories with multi-turn search interactions, leveraging retrieved token masking for stable RL training and a simple outcome-based reward function. Experiments on seven question-answering datasets show that Search-R1 improves performance by 41% (Qwen2.5-7B) and 20% (Qwen2.5-3B) over various RAG baselines under the same setting. This paper further provides empirical insights into RL optimization methods, LLM choices, and response length dynamics in retrieval-augmented reasoning. The code and model checkpoints are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 2.1 Large Language Models and Retrieval
  • 2.2 Large Language Models and Reinforcement Learning
  • 3 Search-R1
  • 3.1 Reinforcement Learning with a Search Engine
  • 3.2 Generation with Multi-turn Search Engine Calling
  • 3.3 Training Template
  • 3.4 Reward Modeling
  • 4 Main Results
  • 4.1 Datasets
  • 4.2 Baselines
  • 4.3 Experimental Setup
  • 4.4 Performance
  • 5 Analysis
  • 5.1 Different RL methods: PPO vs. GRPO
  • 5.2 Base vs. Instruct LLMs
  • 5.3 Response Length and Valid Search Study
  • 5.4 Study of Retrieved Tokens Loss Masking
  • 6 Conclusions
  • References
  • A Formulation of Reinforcement Learning with a Search Engine
  • B Experimental Setups
  • B.1 Baselines
  • B.2 Experimental Settings
  • C Main Results on 14B LLM
  • D Retrieved Token Loss Masking Study
  • E Base vs. Instruct LLMs
  • F Comparison of PPO and GRPO in Search-R1
  • G Number of Retrieved Passages Study in Search-R1 Training
  • H Group Size Study in Search-R1 (GRPO) Training
  • I Comparison between R1 and Search-R1: A Case Study
  • J More Case Studies of Search-R1

Knowls

  1. Knowl 1 — Search-R1 Interleaved Reasoning and Retrieval RL Formulation

    model/method

    SEARCH-R1 extends reinforcement learning (RL) for large language models by framing the search engine as an external environment R\mathcal{R} that interacts with the model during generation. Rather than generating responses solely from the policy model's parametric knowledge πθ(⋅∣x)\pi_\theta(\cdot | x), the policy generates trajectories interleaved with real-time retrieval calls, denoted as y∼πθ(⋅∣x;R)=πθ(⋅∣x)⊗Ry \sim \pi_\theta(\cdot | x; \mathcal{R}) = \pi_\theta(\cdot | x) \otimes \mathcal{R}, where ⊗\otimes represents multi-turn interleaved retrieval and reasoning.

    The policy optimization objective is formulated as:

    max⁡πθEx∼D,y∼πθ(⋅∣x;R)[rϕ(x,y)]−βDKL[πθ(y∣x;R) ∥ πref(y∣x;R)]\max_{\pi_\theta} \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_\theta(\cdot | x; \mathcal{R})} \left[ r_\phi(x, y) \right] - \beta D_{\text{KL}}\left[ \pi_\theta(y | x; \mathcal{R}) \,\Vert\, \pi_{\text{ref}}(y | x; \mathcal{R}) \right]

    where D\mathcal{D} denotes the dataset of input queries xx, yy is the complete generated rollout sequence comprising reasoning steps, search queries, retrieved passages, and final answers, πref\pi_{\text{ref}} is a frozen reference model serving as a regularization anchor, rϕ(x,y)r_\phi(x, y) is the reward function, β\beta is the Kullback-Leibler (KL) divergence regularization weight, and DKLD_{\text{KL}} is the KL-divergence computed over the joint token sequence distribution conditioned on both the input prompt and the retrieved context.

  2. Knowl 2 — Multi-Turn Search and Reasoning Rollout Procedure

    algorithm

    The SEARCH-R1 generation process enables the policy model to alternate between chain-of-thought reasoning and external search engine queries within an action budget BB. The model signals a search by emitting <search> query </search> tokens. When detected, the search engine R\mathcal{R} is queried, and the retrieved content is wrapped in <information> ... </information> tags and appended to the context. If the model outputs <answer> ... </answer>, rollout terminates. If the rollout ends without valid action tags, a rethink prompt is appended.

    Input: Input query xx, policy model πθ\pi_\theta, search engine R\mathcal{R}, maximum action budget BB
    Output: Final response yy
    Initialize rollout sequence y←∅y \leftarrow \emptyset
    Initialize action count b←0b \leftarrow 0
    while b<Bb < B do
        Initialize current action sequence yb←∅y_b \leftarrow \emptyset
        while True do
            Generate response token yt∼πθ(⋅∣x,y+yb)y_t \sim \pi_\theta(\cdot \mid x, y + y_b)
            Append yty_t to rollout sequence: yb←yb+yty_b \leftarrow y_b + y_t
            if yt∈[</search>,</answer>,<eos>]y_t \in [\text{</search>}, \text{</answer>}, \text{<eos>}] then
                break
            end if
        end while
        y←y+yby \leftarrow y + y_b
        if "<search>" and "</search>" detected in yby_b then
            Extract search query q←Parse(yb,<search>,</search>)q \leftarrow \text{Parse}(y_b, \text{<search>}, \text{</search>})
            Retrieve search results d=R(q)d = \mathcal{R}(q)
            Append retrieval: y←y+<information>+d+</information>y \leftarrow y + \text{<information>} + d + \text{</information>}
        else if "<answer>" and "</answer>" detected in yby_b then
            return final generated response yy
        else
            Append rethink message: y←y+"My action is not correct. Let me rethink."y \leftarrow y + \text{"My action is not correct. Let me rethink."}
        end if
        b←b+1b \leftarrow b + 1
    end while
    return final generated response yy
  3. Knowl 3 — PPO Objective with Retrieved Token Loss Masking

    equation

    In SEARCH-R1, rollout trajectories contain both tokens generated by the policy model and external tokens returned by the search engine. To prevent computing gradients on non-generated retrieved tokens, token loss masking I(yt)I(y_t) is applied, where I(yt)=1I(y_t) = 1 if token yty_t was generated by the language model and I(yt)=0I(y_t) = 0 if yty_t is part of retrieved context.

    The Proximal Policy Optimization (PPO) training objective with search and token masking is:

    JPPO(θ)=Ex∼D,y∼πold(⋅∣x;R)[1∑t=1∣y∣I(yt)∑t=1:I(yt)=1∣y∣min⁡(πθ(yt∣x,y<t;R)πold(yt∣x,y<t;R)At,clip(πθ(yt∣x,y<t;R)πold(yt∣x,y<t;R),1−ϵ,1+ϵ)At)]J_{\text{PPO}}(\theta) = \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_{\text{old}}(\cdot | x; \mathcal{R})} \left[ \frac{1}{\sum_{t=1}^{|y|} I(y_t)} \sum_{t=1: I(y_t)=1}^{|y|} \min \left( \frac{\pi_\theta(y_t | x, y_{<t}; \mathcal{R})}{\pi_{\text{old}}(y_t | x, y_{<t}; \mathcal{R})} A_t, \text{clip}\left( \frac{\pi_\theta(y_t | x, y_{<t}; \mathcal{R})}{\pi_{\text{old}}(y_t | x, y_{<t}; \mathcal{R})}, 1 - \epsilon, 1 + \epsilon \right) A_t \right) \right]

    where πθ\pi_\theta is the current policy model, πold\pi_{\text{old}} is the behavior policy model from the previous rollout step, ϵ\epsilon is the clipping hyperparameter (set to 0.20.2), and AtA_t is the advantage estimate computed via Generalized Advantage Estimation (GAE with λ=1,γ=1\lambda = 1, \gamma = 1) utilizing future rewards {r≥t}\{r_{\ge t}\} and a learned value function VϕV_\phi.

  4. Knowl 4 — GRPO Objective with Retrieved Token Loss Masking

    equation

    Group Relative Policy Optimization (GRPO) eliminates the need for a separate critic/value network by estimating baseline rewards from a group of GG rollouts sampled for each input prompt xx.

    The GRPO objective with search engine rollouts and retrieved token masking is:

    JGRPO(θ)=Ex∼D,{yi}i=1G∼πold(⋅∣x;R)[1G∑i=1G1∑t=1∣yi∣I(yi,t)∑t=1:I(yi,t)=1∣yi∣min⁡(πθ(yi,t∣x,yi,<t;R)πold(yi,t∣x,yi,<t;R)A^i,t,clip(πθ(yi,t∣x,yi,<t;R)πold(yi,t∣x,yi,<t;R),1−ϵ,1+ϵ)A^i,t)−βDKL[πθ ∥ πref]]J_{\text{GRPO}}(\theta) = \mathbb{E}_{x \sim \mathcal{D}, \{y_i\}_{i=1}^G \sim \pi_{\text{old}}(\cdot | x; \mathcal{R})} \left[ \frac{1}{G} \sum_{i=1}^G \frac{1}{\sum_{t=1}^{|y_i|} I(y_{i,t})} \sum_{t=1: I(y_{i,t})=1}^{|y_i|} \min \left( \frac{\pi_\theta(y_{i,t} | x, y_{i,<t}; \mathcal{R})}{\pi_{\text{old}}(y_{i,t} | x, y_{i,<t}; \mathcal{R})} \hat{A}_{i,t}, \text{clip}\left( \frac{\pi_\theta(y_{i,t} | x, y_{i,<t}; \mathcal{R})}{\pi_{\text{old}}(y_{i,t} | x, y_{i,<t}; \mathcal{R})}, 1 - \epsilon, 1 + \epsilon \right) \hat{A}_{i,t} \right) - \beta D_{\text{KL}}\left[ \pi_\theta \,\Vert\, \pi_{\text{ref}} \right] \right]

    where I(yi,t)I(y_{i,t}) is the binary token mask indicator (11 for model-generated tokens, 00 for retrieved passage tokens), ϵ\epsilon is the clipping ratio (0.20.2), β\beta is the KL regularization coefficient (0.0010.001), and A^i,t\hat{A}_{i,t} is the standardized relative advantage of response yiy_i computed from the outcome rewards within the sampled group {y1,…,yG}\{y_1, \dots, y_G\}. Retrieved token masking is also applied when computing the KL divergence loss DKLD_{\text{KL}}.

  5. Knowl 5 — Search-R1 Structural Training Template and Outcome Reward

    model/method

    SEARCH-R1 guides the policy model using a minimal structural system template without enforcing explicit step-by-step heuristics or content-specific biases:

    "Answer the given question. You must conduct reasoning inside <think> and </think> first every time you get new information. After reasoning, if you find you lack some knowledge, you can call a search engine by <search> query </search>, and it will return the top searched results between <information> and </information>. You can search as many times as you want. If you find no further external knowledge needed, you can directly provide the answer inside <answer> and </answer> without detailed illustrations. For example, <answer> xxx </answer>. Question: {question}."

    The reward signal is strictly rule-based and outcome-driven, using binary Exact Match (EM) between the extracted answer apreda_{\text{pred}} and ground truth answer agolda_{\text{gold}}:

    rϕ(x,y)=EM(apred,agold)r_\phi(x, y) = \text{EM}(a_{\text{pred}}, a_{\text{gold}})

    No neural reward models or intermediate step/format rewards are used during training, avoiding the complexity and instability of process-based supervision.

  6. Knowl 6 — Search-R1 Question Answering Performance Across Benchmarks

    data/table

    SEARCH-R1 was evaluated on seven question answering benchmarks across general QA (Natural Questions [NQ], TriviaQA, PopQA) and multi-hop QA (HotpotQA, 2WikiMultiHopQA [2wiki], Musique, Bamboogle) using Exact Match (EM). Training was performed on a combined dataset of NQ and HotpotQA training splits, using the 2018 Wikipedia dump with the E5 dense retriever (top-33 passages). Evaluation covers both in-domain (†^\dagger) and out-of-domain (⋆^\star) benchmarks.

    Method NQ†^\dagger TriviaQA⋆^\star PopQA⋆^\star HotpotQA†^\dagger 2wiki⋆^\star Musique⋆^\star Bamboogle⋆^\star Avg.
    Qwen2.5-7B (Base / Instruct)
    Direct Inference 0.134 0.408 0.140 0.183 0.250 0.031 0.120 0.181
    CoT 0.048 0.185 0.054 0.092 0.111 0.022 0.232 0.106
    IRCoT 0.224 0.478 0.301 0.133 0.149 0.072 0.224 0.239
    Search-o1 0.151 0.443 0.131 0.187 0.176 0.058 0.296 0.206
    RAG 0.349 0.585 0.392 0.299 0.235 0.058 0.208 0.304
    SFT 0.318 0.354 0.121 0.217 0.259 0.066 0.112 0.207
    R1-base 0.297 0.539 0.202 0.242 0.273 0.083 0.296 0.276
    R1-instruct 0.270 0.537 0.199 0.237 0.292 0.072 0.293 0.271
    Rejection Sampling 0.360 0.592 0.380 0.331 0.296 0.123 0.355 0.348
    SEARCH-R1-base 0.480 0.638 0.457 0.433 0.382 0.196 0.432 0.431
    SEARCH-R1-instruct 0.393 0.610 0.397 0.370 0.414 0.146 0.368 0.385
    Qwen2.5-3B (Base / Instruct)
    Direct Inference 0.106 0.288 0.108 0.149 0.244 0.020 0.024 0.134
    CoT 0.023 0.032 0.005 0.021 0.021 0.002 0.000 0.015
    IRCoT 0.111 0.312 0.200 0.164 0.171 0.067 0.240 0.181
    Search-o1 0.238 0.472 0.262 0.221 0.218 0.054 0.320 0.255
    RAG 0.348 0.544 0.387 0.255 0.226 0.047 0.080 0.270
    SFT 0.249 0.292 0.104 0.186 0.248 0.044 0.112 0.176
    R1-base 0.226 0.455 0.173 0.201 0.268 0.055 0.224 0.229
    R1-instruct 0.210 0.449 0.171 0.208 0.275 0.060 0.192 0.224
    Rejection Sampling 0.294 0.488 0.332 0.240 0.233 0.059 0.210 0.265
    SEARCH-R1-base 0.406 0.587 0.435 0.284 0.273 0.049 0.088 0.303
    SEARCH-R1-instruct 0.341 0.545 0.378 0.324 0.319 0.103 0.264 0.325

    SEARCH-R1 achieves an average relative improvement of 41% (Qwen2.5-7B) and 20% (Qwen2.5-3B) over standard RAG baselines, and also outperforms parametric reasoning RL (R1) across all evaluated benchmarks.

  7. Knowl 7 — Ablation of Retrieved Token Loss Masking

    data/table

    Masking retrieved passage tokens during policy gradient updates is critical for stable learning. When loss masking is omitted (i.e., gradients are computed over both generated and retrieved tokens), performance degrades substantially across all evaluated QA datasets.

    Method NQ TriviaQA PopQA HotpotQA 2wiki Musique Bamboogle Avg.
    Qwen2.5-7B-Base (PPO)
    SEARCH-R1 w. mask 0.480 0.638 0.457 0.433 0.382 0.196 0.432 0.431
    SEARCH-R1 w.o. mask 0.388 0.567 0.391 0.325 0.321 0.108 0.304 0.343
    Qwen2.5-3B-Base (PPO)
    SEARCH-R1 w. mask 0.406 0.587 0.435 0.284 0.273 0.049 0.088 0.303
    SEARCH-R1 w.o. mask 0.346 0.484 0.365 0.241 0.244 0.053 0.104 0.262

    On Qwen2.5-7B-base, retrieved token masking improves overall average Exact Match from 0.343 to 0.431 (+25.7% relative gain), preventing erratic gradient signals caused by external non-generated text.

  8. Knowl 8 — Comparison of PPO and GRPO in Retrieval-Augmented Reasoning

    empirical result

    When comparing Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO) for SEARCH-R1 on Qwen2.5 models:

    1. Convergence Speed: GRPO converges faster in initial training steps across both base and instruction-tuned models because it does not require a critic model warm-up phase.
    2. Training Stability: PPO provides superior optimization stability over extended training (500 steps). GRPO exhibits susceptibility to reward collapse after many training steps, especially with larger group sizes (e.g., group size G=5G = 5).
    3. Final Performance: Both optimization methods achieve comparable peak training rewards and test performance (e.g., Qwen2.5-7B-instruct achieves 0.396 average EM with GRPO vs. 0.385 with PPO; Qwen2.5-7B-base achieves 0.431 average EM with PPO vs. 0.350 with GRPO at step 500 due to GRPO late-stage divergence).
  9. Knowl 9 — Effect of Retrieved Passage Count on RL Optimization

    data/table

    The number of retrieved passages (top-kk) returned per search call impacts both RL training stability and downstream QA performance during SEARCH-R1 training on Qwen2.5-7B-Base with PPO:

    Setting NQ TriviaQA PopQA HotpotQA 2wiki Musique Bamboogle Avg.
    top-k=1k = 1 0.426 0.614 0.422 0.393 0.296 0.146 0.328 0.375
    top-k=3k = 3 0.480 0.638 0.457 0.433 0.382 0.196 0.432 0.431
    top-k=5k = 5 0.479 0.634 0.440 0.394 0.343 0.156 0.352 0.400

    Setting top-k=3k = 3 yields the highest reward and overall accuracy (0.431 average EM). A setting of top-k=1k = 1 suffers from insufficient recall (0.375 average EM), while top-k=5k = 5 initially converges rapidly within 200 steps but degrades and destabilizes later in training due to noisy, irrelevant passages that lower context precision.

  10. Knowl 10 — Training Dynamics of Response Length, Search Frequency, and Model Pretraining

    empirical result

    Analysis of SEARCH-R1 training dynamics reveals characteristic behavioral patterns:

    1. Response Length Evolution: Trajectory length follows a distinct decrease-increase-stabilize trajectory. In the early training phase (steps 0–100), response length drops sharply while training reward increases slightly as the model learns to discard excessive conversational filler tokens. In the later phase (>100 steps), response length increases significantly alongside reward because the model learns to issue search queries frequently and ingest the resulting multi-paragraph context.
    2. Search Frequency: The number of valid search engine invocations increases monotonically throughout training as the policy discovers that querying external information correlates with positive outcome rewards.
    3. Base vs. Instruction-Tuned Models: Instruction-tuned models begin with higher initial rewards and converge faster. However, base models subjected to RL directly catch up over extended training, achieving comparable or superior final performance (e.g., Qwen2.5-7B-base reaching 0.431 average EM vs. 0.385 for Qwen2.5-7B-instruct).

Coverage note — Omitted specific full-text case study transcripts from Appendix I/J, hardware environment details, and the 14B ablation table which mirrors the 7B findings.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Olivier Pietquin, Ahmet Üstün, and Sara Hooker. Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms. arXiv preprint arXiv:2402.14740, 2024.
  3. 3.Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. Large language models for mathematical reasoning: Progresses and challenges. arXiv preprint arXiv:2402.00157, 2024.
  4. 4.Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. Self-rag: Learning to retrieve, generate, and critique through self-reflection. 2024.
  5. 5.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. Journal of Machine Learning Research, 25(70):1–53, 2024.
  6. 6.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  7. 7.Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997, 2, 2023.
  8. 8.Michael Glass, Gaetano Rossiello, Md Faisal Mahbub Chowdhury, Ankita Rajaram Naik, Pengshan Cai, and Alfio Gliozzo. Re2g: Retrieve, rerank, generate. arXiv preprint arXiv:2207.06300, 2022.
  9. 9.Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, et al. Deepseek-coder: When the large language model meets programming–the rise of code intelligence. arXiv preprint arXiv:2401.14196, 2024.
  10. 10.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025.
  11. 11.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  12. 12.Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps. arXiv preprint arXiv:2011.01060, 2020.
  13. 13.Zhenyu Hou, Xin Lv, Rui Lu, Jiajie Zhang, Yujiang Li, Zijun Yao, Juanzi Li, Jie Tang, and Yuxiao Dong. Advancing language model reasoning through reinforcement learning and inference scaling. arXiv preprint arXiv:2501.11651, 2025.
  14. 14.Sheryl Hsu, Omar Khattab, Chelsea Finn, and Archit Sharma. Grounding by trying: Llms with reinforcement learning-enhanced retrieval. arXiv preprint arXiv:2410.23214, 2024.
  15. 15.Jie Huang and Kevin Chen-Chuan Chang. Towards reasoning in large language models: A survey. arXiv preprint arXiv:2212.10403, 2022.
  16. 16.Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024.
  17. 17.Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. Active retrieval augmented generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 7969–7992, 2023.
  18. 18.Bowen Jin, Jinsung Yoon, Jiawei Han, and Sercan O Arik. Long-context llms meet rag: Overcoming challenges for long inputs in rag. In The Thirteenth International Conference on Learning Representations, 2024.
  19. 19.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551, 2017.
  20. 20.Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore. Reinforcement learning: A survey. Journal of artificial intelligence research, 4:237–285, 1996.
  21. 21.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick SH Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. Dense passage retrieval for open-domain question answering. In EMNLP (1), pp. 6769–6781, 2020.
  22. 22.Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier. A survey of reinforcement learning from human feedback. arXiv preprint arXiv:2312.14925, 10, 2023.
  23. 23.Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, et al. Training language models to self-correct via reinforcement learning. arXiv preprint arXiv:2409.12917, 2024.
  24. 24.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466, 2019.
  25. 25.Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, et al. Rewardbench: Evaluating reward models for language modeling. arXiv preprint arXiv:2403.13787, 2024.
  26. 26.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33:9459–9474, 2020.
  27. 27.Xiaoxi Li, Jiajie Jin, Yujia Zhou, Yongkang Wu, Zhonghua Li, Qi Ye, and Zhicheng Dou. Retrollm: Empowering large language models to retrieve fine-grained evidence within generation. arXiv preprint arXiv:2412.11919, 2024.
  28. 28.Xiaoxi Li, Guanting Dong, Jiajie Jin, Yuyao Zhang, Yujia Zhou, Yutao Zhu, Peitian Zhang, and Zhicheng Dou. Search-o1: Agentic search-enhanced large reasoning models. arXiv preprint arXiv:2501.05366, 2025.
  29. 29.Yinheng Li, Shaofei Wang, Han Ding, and Hang Chen. Large language models in finance: A survey. In Proceedings of the fourth ACM international conference on AI in finance, pp. 374–382, 2023.
  30. 30.Xi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi, Maria Lomeli, Richard James, Pedro Rodriguez, Jacob Kahn, Gergely Szilvasy, Mike Lewis, et al. Ra-dit: Retrieval-augmented dual instruction tuning. In The Twelfth International Conference on Learning Representations, 2023.
  31. 31.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, and Daniel Khashabi. When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories. arXiv preprint arXiv:2212.10511, 7, 2022.
  32. 32.Yu Meng, Mengzhou Xia, and Danqi Chen. Simpo: Simple preference optimization with a reference-free reward. Advances in Neural Information Processing Systems, 37:124198–124235, 2024.
  33. 33.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744, 2022.
  34. 34.Richard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho, Sainbayar Sukhbaatar, and Jason Weston. Iterative reasoning preference optimization. Advances in Neural Information Processing Systems, 37:116617–116637, 2024.
  35. 35.Cheng Peng, Xi Yang, Aokun Chen, Kaleb E Smith, Nima PourNejatian, Anthony B Costa, Cheryl Martin, Mona G Flores, Ying Zhang, Tanja Magoc, et al. A study of generative large language model for medical research and healthcare. NPJ digital medicine, 6(1):210, 2023.
  36. 36.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. Measuring and narrowing the compositionality gap in language models. arXiv preprint arXiv:2210.03350, 2022.
  37. 37.Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-Rong Wen. Tool learning with large language models: A survey. Frontiers of Computer Science, 19(8):198343, 2025.
  38. 38.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36:53728–53741, 2023.
  39. 39.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36: 68539–68551, 2023.
  40. 40.John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438, 2015.
  41. 41.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  42. 42.Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.
  43. 43.Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. Hybridflow: A flexible and efficient rlhf framework. arXiv preprint arXiv:2409.19256, 2024.
  44. 44.Richard S Sutton, Andrew G Barto, et al. Reinforcement learning. Journal of Cognitive Neuroscience, 11(1):126–134, 1999.
  45. 45.Gemini Team. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024.
  46. 46.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. arXiv preprint arXiv:2212.10509, 2022a.
  47. 47.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. Musique: Multihop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10:539–554, 2022b.
  48. 48.Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533, 2022.
  49. 49.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022.
  50. 50.Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun, Kang Liu, and Jun Zhao. Large language models are better reasoners with self-verification. arXiv preprint arXiv:2212.09561, 2022.
  51. 51.Ronald J Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8:229–256, 1992.
  52. 52.Tian Xie, Zitian Gao, Qingnan Ren, Haoming Luo, Yuqian Hong, Bryan Dai, Joey Zhou, Kai Qiu, Zhirong Wu, and Chong Luo. Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning. arXiv preprint arXiv:2502.14768, 2025.
  53. 53.Guangzhi Xiong, Qiao Jin, Xiao Wang, Yin Fang, Haolin Liu, Yifan Yang, Fangyuan Chen, Zhixing Song, Dengyu Wang, Minjia Zhang, et al. Rag-gym: Optimizing reasoning and search agents with process supervision. arXiv preprint arXiv:2502.13957, 2025.
  54. 54.An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115, 2024.
  55. 55.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600, 2018.
  56. 56.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), 2023.
  57. 57.Yue Yu, Wei Ping, Zihan Liu, Boxin Wang, Jiaxuan You, Chao Zhang, Mohammad Shoeybi, and Bryan Catanzaro. Rankrag: Unifying context ranking with retrieval-augmented generation in llms. Advances in Neural Information Processing Systems, 37:121156–121184, 2024.
  58. 58.Zhenrui Yue, Honglei Zhuang, Aijun Bai, Kai Hui, Rolf Jagerman, Hansi Zeng, Zhen Qin, Dong Wang, Xuanhui Wang, and Michael Bendersky. Inference scaling for long-context retrieval augmented generation. arXiv preprint arXiv:2410.04343, 2024.
  59. 59.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. Siren’s song in the ai ocean: a survey on hallucination in large language models. arXiv preprint arXiv:2309.01219, 2023.
  60. 60.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 1(2), 2023.
  61. 61.Wayne Xin Zhao, Jing Liu, Ruiyang Ren, and Ji-Rong Wen. Dense text retrieval based on pretrained language models: A survey. ACM Transactions on Information Systems, 42(4): 1–60, 2024.

Citation

MLA
Jin, B., et al. “Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning”. arXiv, 2025, http://arxiv.org/abs/2503.09516v5.
APA
Jin, B., Zeng, H., Yue, Z., Yoon, J., Arik, S., Wang, D., Zamani, H., & Han, J. (2025). Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning. arXiv. http://arxiv.org/abs/2503.09516v5
Chicago
Jin, B., H. Zeng, Z. Yue, et al. 2025. “Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning”. arXiv. http://arxiv.org/abs/2503.09516v5.
Harvard
Jin, B. et al. (2025) “Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2503.09516v5.
Vancouver
1. Jin B, Zeng H, Yue Z, Yoon J, Arik S, Wang D, Zamani H, Han J (2025) Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning. arXiv

BibTeX

@article{jin2025search,
  title = {Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning},
  author = {Jin, Bowen and Zeng, Hansi and Yue, Zhenrui and Yoon, Jinsung and Arik, Sercan and Wang, Dong and Zamani, Hamed and Han, Jiawei},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2503.09516v5},
  eprint = {2503.09516}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission