Deep Researcher with Test-time Diffusion

Rujun HanYanfei ChenGuan SunLesly MiculicichZoey CuiZhuYuanjun (Sophia) BiWeiming WenHui WanChunfeng WenSolène Maître

article2025arXiv28 citations

Proposes Test-Time Diffusion Deep Researcher, a framework that models long-form report generation as an iterative diffusion process to refine structured drafts through continuous retrieval, outperforming existing LLM research agents on complex multi-hop reasoning benchmarks.

Listen

Autonomous research systems powered by large language models have advanced rapidly, yet their ability to generate complex, long-form research reports frequently hits a ceiling. Most existing systems operate through rigid, linear, or disconnected search workflows that struggle to maintain broad context, resulting in significant information loss and uncoordinated findings. The article introduces and evaluates Test-Time Diffusion Deep Researcher (TTD-DR), a framework designed to mimic iterative human writing behavior by framing report generation as a diffusion process that continuously refines an initial draft with external web search and component-level optimization.

The framework operates across three key stages: generating a research plan, conducting iterative information retrieval, and synthesizing a final report. Credibility and robustness are established through two core mechanisms. First, report-level denoising uses the evolving draft to dynamically guide search queries while immediately integrating retrieved findings back into the text across up to 20 revision steps. Second, component-wise self-evolution optimizes individual sub-tasks—such as planning, query generation, and answer synthesis—by generating diverse variants, scoring them with an automated evaluator, refining them based on critique, and merging the best elements. The system was evaluated across long-form report benchmarks (LongForm Research and DeepConsult) and complex multi-hop reasoning datasets (Humanity's Last Exam text subsets and GAIA), benchmarked against proprietary and open-source systems including OpenAI Deep Research, Perplexity Deep Research, and Grok DeeperSearch.

The evaluation demonstrates several key findings:

  1. TTD-DR achieves state-of-the-art performance in long-form report generation, recording a 69.1% win rate against OpenAI Deep Research on the LongForm Research benchmark and a 74.5% win rate on DeepConsult.
  2. In multi-hop reasoning and concise question answering, TTD-DR outperforms OpenAI Deep Research by 4.8 percentage points on search-intensive academic queries (33.9% vs. 29.1%) and by 7.7 percentage points across the broader academic benchmark (34.3% vs. 26.6%), while also leading on real-world general tasks (69.1% vs. 67.4%).
  3. Denoising with retrieval increases search query novelty by more than 12 percentage points and captures over 51% of the final report's key information by the ninth search step, outperforming 20 steps of self-evolution alone.
  4. Latency-efficiency analyses show that TTD-DR exhibits a steeper efficiency frontier than baseline methods, yielding greater output quality per unit of processing time.

These findings indicate that adopting human-like drafting and continuous revision loops fundamentally improves coherence, depth, and factual retention in automated research. Unlike conventional systems that perform isolated searches before writing, maintaining a central working draft reduces information loss and prevents circular or redundant queries. For organizations seeking automated strategic analysis, market intelligence, or technical synthesis, this architecture provides a viable pathway to higher-quality outputs using standard search engines without relying on opaque, proprietary toolstacks.

Organizations developing or deploying automated research agents should consider transitioning from linear multi-agent pipelines to draft-centric, iterative revision frameworks. Engineering teams should prioritize test-time scaling methods that combine sub-task self-evolution with dynamic retrieval loops. However, decision-makers should note that this study focused solely on text-based web search; the current framework does not integrate code execution or multimodal web browsing environments. Additionally, because evaluations rely on automated model judges calibrated to human preferences, critical strategic deployments should maintain human-in-the-loop oversight until further real-world pilot validations are completed.

Cover for Deep Researcher with Test-time Diffusion

Abstract

Deep research agents, powered by Large Language Models (LLMs), are rapidly advancing; yet, their performance often plateaus when generating complex, long-form research reports using generic test-time scaling algorithms. Drawing inspiration from the iterative nature of human research, which involves cycles of searching, reasoning, and revision, we propose the Test-Time Diffusion Deep Researcher (TTD-DR). This novel framework conceptualizes research report generation as a diffusion process. TTD-DR initiates this process with a preliminary draft, an updatable skeleton that serves as an evolving foundation to guide the research direction. The draft is then iteratively refined through a "denoising" process, which is dynamically informed by a retrieval mechanism that incorporates external information at each step. The core process is further enhanced by a self-evolutionary algorithm applied to each component of the agentic workflow, ensuring the generation of high-quality context for the diffusion process. This draft-centric design makes the report writing process more timely and coherent while reducing information loss during the iterative search process. We demonstrate that our TTD-DR achieves state-of-the-art results on a wide array of benchmarks that require intensive search and multi-hop reasoning, significantly outperforming existing deep research agents.

Table of Contents

  • 1 Introduction
  • 2 Test-Time Diffusion Deep Researcher (TTD-DR)
  • 2.1 Backbone Deep Research Agent
  • 2.2 Component-wise Self-Evolution
  • 2.3 Report-level Denoising with Retrieval
  • 3 Experimental Setup
  • 3.1 Evaluation Metrics
  • 3.2 LLM-as-a-judge Calibration
  • 3.3 Data
  • 3.4 Implementation Details
  • 3.5 Compared Systems
  • 4 Results and Analysis
  • 4.1 Main Results
  • 4.2 Analysis
  • 5 Related Work
  • 6 Conclusions
  • References
  • A Appendix
  • A.1 Evaluation Guidelines
  • A.2 Human Annotation Interface
  • A.3 Human and LLM-as-a-judge Alignment
  • A.4 HLE Query Categorization
  • A.5 Answer Merging.
  • A.6 Hyper-parameters
  • A.7 Question Complexity
  • A.8 Answer Complexity
  • A.9 Query Novelty
  • A.10 Report Coverage
  • A.11 Additional Analysis Results

Knowls

  1. Knowl 1 — Draft-driven test-time diffusion for deep research

    model/method

    Test-Time Diffusion Deep Researcher (TTD-DR) treats long-form research-report generation as an iterative refinement process. Given a user query, the system creates an initial draft report and a research plan. The draft is an editable global scaffold: it guides what to search for next, while retrieved information is used to verify or revise the draft. This couples the evolving report to the research trajectory instead of postponing report synthesis until after independent searches are complete. Component-wise self-evolution improves the quality of selected workflow outputs, supplying stronger context to the draft-refinement loop. The intended effect is more coherent integration of findings and less information loss during multi-step research.

  2. Knowl 2 — Three-stage backbone research agent

    model/method

    The TTD-DR backbone processes a query through three stages. First, a planning agent creates a structured outline of key areas needed in the final report. Second, an iterative search-and-synthesis workflow generates questions using the user query, plan, current context, and earlier question-answer pairs; a search system retrieves relevant documents and synthesizes them into answers rather than merely accumulating raw documents. Search continues until the plan is adequately covered or the iteration limit is reached. Third, a report-generation agent synthesizes the plan and the complete set of search question-answer pairs into a coherent final report. In TTD-DR, the evolving draft is also supplied to question generation and is revised using each new synthesized answer.

  3. Knowl 3 — Retrieval-guided report denoising

    algorithm

    This procedure refines an initial report draft by alternating targeted search with draft revision. For the reported implementation, the maximum number of denoising and retrieval steps is 20. The question generator uses the user query, research plan, current draft, and prior questions and answers; the answer-search component retrieves external information and returns a synthesized answer; the revision component uses the query, previous draft, and accumulated search history to update the draft.

    Input: user query qq, research plan PP, initial draft R0R_0, empty question history QQ, empty answer history AA, maximum steps N=20N=20
    Output: final report
    Set current draft RR to R0R_0
    For t=1,…,Nt=1,\ldots,N:
        Generate question QtQ_t using qq, PP, RR, and histories Q,AQ,A; target gaps in the draft or plan
        Add QtQ_t to QQ
        Retrieve and synthesize answer AtA_t for QtQ_t
        Add AtA_t to AA
        Revise RR using qq, the previous draft, and updated histories Q,AQ,A
        If the research is judged complete or the workflow signals exit, stop the loop
    Generate the final report from PP, all search answers, and the accumulated draft revisions
    Return the final report

    The preliminary draft is generated from the query before this loop. The draft after each revision guides the next search question; the final report is generated after the loop from the plan and the full search and revision history.

  4. Knowl 4 — Component-wise self-evolution of workflow agents

    algorithm

    TTD-DR improves workflow components by sampling multiple candidate outputs, evaluating them with an LLM judge, revising candidates using the judge's scores and textual critiques, and merging the resulting candidates into one output. The method is applicable to plan, search-question, answer, and report generation. For search answers, the candidates are alternative answers conditioned on earlier workflow context; the judge assesses criteria such as helpfulness and comprehensiveness. Each revision episode repeats evaluation and revision up to its configured number of steps, after which a crossover or merge prompt combines information from the candidate paths.

    The reported settings are: for LongForm Research and DeepConsult, initial candidate counts are 1 plan, 5 questions, 3 answers, and 1 report; evolution-step counts are 1 for plans, 0 for questions, 0 for answers, and 1 for reports. For HLE and GAIA, initial counts are 1 plan, 5 questions, 3 answers, and 5 reports; evolution-step counts are 1 for plans, 0 for questions, 0 for answers, and 0 for reports. Thus, multiple candidates do not necessarily undergo iterative revision: the configured number of evolution steps determines whether the feedback-and-revision loop is used.

  5. Knowl 5 — Evaluation design and implementation

    experimental setup

    TTD-DR is implemented using Gemini-2.5-pro as its base model, the Google Agent Development Kit to orchestrate agents and workflows, and grounding with Google Search for retrieval. The retrieval-and-denoising loop is capped at 20 steps. Evaluation covers long-form research reports and search-intensive multi-hop questions: LongForm Research contains 205 licensed real-world queries across industry domains; DeepConsult contains business and consulting prompts; HLE-full is the text-only set of 2,500 Humanity's Last Exam questions; HLE-search is a random sample of 200 HLE questions classified by Gemini-1.5-pro as requiring search; and GAIA is the benchmark evaluation set.

    For long-form reports, raters compare pairs using helpfulness (user-intent satisfaction, clarity and coherence, accuracy, and appropriate language) and comprehensiveness (whether key information is missing). Human pairwise judgments are used to calibrate an LLM judge: 200 report comparisons involving TTD-DR and OpenAI Deep Research were assessed, with each pair rated twice by humans. Gemini-1.5-pro is used as the final long-form evaluator. For HLE and GAIA correctness, the evaluator extracts a short answer and compares it with the ground truth using the standard evaluation prompt; this correctness judge is not human-calibrated. The evaluation focuses on search and does not add browsing or coding tools.

  6. Knowl 6 — Benchmark comparison with leading research agents

    data/table

    The comparison evaluates long-form report quality by win rate against OpenAI Deep Research and short-form benchmark performance by correctness. TTD-DR has the highest reported score on every listed benchmark: its win rates are 69.1% on LongForm Research and 74.5% on DeepConsult, and its correctness is 33.9% on HLE-search, 34.3% on HLE-full, and 69.1% on GAIA. OpenAI Deep Research is the reference for win rates; its entries for those two columns are therefore not reported. The table also shows that TTD-DR exceeds OpenAI Deep Research correctness by 4.8 percentage points on HLE-search, 7.7 points on HLE-full, and 1.7 points on GAIA.

    System LongForm Research win rate (%) DeepConsult win rate (%) HLE-search correctness (%) HLE-full correctness (%) GAIA correctness (%)
    OpenAI Deep Research – – 29.1 26.6 67.4
    Perplexity Deep Research 21.8 32.0 14.5 21.1 54.5
    Grok DeeperSearch 16.1 16.0 19.3 – 47.9
    GPT-Researcher 18.3 9.4 2.0 4.1 37.7
    Open Deep Search 2.6 2.2 3.0 0.4 20.9
    TTD-DR 69.1 74.5 33.9 34.3 69.1

    Win rate is a side-by-side comparison against OpenAI Deep Research; correctness is the percentage of answers judged to match the benchmark reference. Grok DeeperSearch has no reported HLE-full value.

  7. Knowl 7 — Ablations show gains from both proposed mechanisms

    data/table

    The ablation compares standalone Gemini models, those models with a search tool, the TTD-DR backbone, and successive additions of self-evolution and retrieval-guided diffusion. All win rates are against OpenAI Deep Research; correctness is benchmark answer accuracy. Adding self-evolution to the backbone raises the LongForm Research win rate from 39.4% to 60.9% and DeepConsult from 24.5% to 59.8%. Adding denoising with retrieval then raises these to 69.1% and 74.5%, respectively, and produces the best listed score on all three correctness benchmarks. The results support contributions from both test-time mechanisms, with the final system outperforming its backbone and the search-augmented standalone LLM baselines.

    System LongForm Research win rate (%) DeepConsult win rate (%) HLE-search correctness (%) HLE-full correctness (%) GAIA correctness (%)
    OpenAI Deep Research – – 29.1 26.6 67.4
    Gemini-2.5-flash 21.0 16.7 2.8 11.6 31.5
    Gemini-2.5-flash with search 27.8 17.6 14.6 14.6 57.6
    Gemini-2.5-pro 31.0 17.6 8.6 20.9 57.0
    Gemini-2.5-pro with search 35.0 19.6 20.0 21.6 61.8
    Backbone DR Agent 39.4 24.5 26.8 28.6 61.8
    Backbone DR Agent with self-evolution 60.9 59.8 30.6 29.4 63.0
    TTD-DR with denoising and retrieval 69.1 74.5 33.9 34.3 69.1

    The standalone LLM rows use Gemini-2.5-flash or Gemini-2.5-pro, with and without a simple search tool. The TTD-DR variants use Gemini-2.5-pro as their base model.

  8. Knowl 8 — Process analyses indicate complementary component effects

    empirical result

    The authors analyze generated search content and information flow to characterize how the two mechanisms help. On DeepConsult, self-evolution increases the cumulative number of distinct key points extracted from search questions and answers relative to the backbone agent, indicating richer search questions and synthesized answers. Compared with self-evolution, denoising with retrieval increases cumulative novel search-question points by more than 12 percentage points across the search-and-revision process. By step 9, denoising with retrieval has incorporated 51.2% of the information attributed to the final report in search answers; at that stage it exceeds the win ratio of self-evolution run for 20 search steps by 4.2 percentage points. These measurements support the authors' interpretation that self-evolution enriches information and that draft-guided retrieval surfaces and preserves useful information earlier.

  9. Knowl 9 — Performance gains with additional test-time compute

    empirical result

    The authors compare performance with latency as retrieval-and-revision steps are increased up to 20. On LongForm Research, adding steps improves TTD-DR's win rate against OpenAI Deep Research; the reported performance-latency frontier places the final system above or on par with competing research agents at similar latency. Across the tested designs—Gemini-2.5-pro with search, the backbone agent, the backbone with self-evolution, and TTD-DR with denoising and retrieval—the final design has the steepest performance-versus-latency slope. The authors interpret this as evidence that both self-evolution and denoising with retrieval are effective test-time scaling methods. The same analysis on HLE-search shows a similar ordering, while the authors note that its short-answer task is less aligned with their primary long-form-report use case.

  10. Knowl 10 — Scope limitations

    limitation

    The evaluated TTD-DR system focuses on search-tool use and does not incorporate other tools such as web browsing or coding. The work also focuses on test-time scaling rather than training or tuning the research agents; agent tuning is left for future work.

Coverage note — The detailed annotation rubrics, prompt templates, and per-domain query-distribution percentages are omitted because they are evaluation implementation details rather than central methods or findings; the major contributed methods, evaluations, results, analyses, and stated limitations are represented.

References

  1. 1.S. Alzubi, C. Brooks, P. Chiniya, E. Contente, C. von Gerlach, L. Irwin, Y. Jiang, A. Kaz, W. Nguyen, S. Oh, H. Tyagi, and P. Viswanath. Open deep search: Democratizing search with open-source reasoning agents. 03 2025. URL https://arxiv.org/abs/2503.20201.
  2. 2.J. Baek, S. K. Jauhar, S. Cucerzan, and S. J. Hwang. Researchagent: Iterative research idea generation over scientific literature with large language models. April 2024.
  3. 3.A. Catalano. Patterns of graduate students’ information seeking behavior: a meta-synthesis of the literature", journal of documentation. Patterns of graduate students’ information seeking behavior: a meta-synthesis of the literature, 69(2):243–274, 2013. URL https://doi.org/10.1108/00220411311300066.
  4. 4.Q. Chen, M. Yang, L. Qin, J. Liu, Z. Yan, J. Guan, D. Peng, Y. Ji, H. Li, M. Hu, Y. Zhang, Y. Liang, Y. Zhou, J. Wang, Z. Chen, and W. Che. Ai4research: A survey of artificial intelligence for scientific research. 07 2025. URL https://arxiv.org/pdf/2507.01903.
  5. 5.M. S. Chitwood. Do you know the steps of the writing process?, 2022. URL https:https://melanieschitwood.com/do-you-know-the-steps-of-the-writing-process/.
  6. 6.J. Coelho, J. Ning, J. He, K. Mao, A. Paladugu, P. Setlur, J. Jin, J. Callan, J. Magalhães, B. Martins, and C. Xiong. Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research. 05 2025. doi: 10.48550/arXiv.2505.19253.
  7. 7.DeerFlow. Deerflow, 2025. URL https://github.com/bytedance/deer-flow.
  8. 8.L. Flower and J. R. Hayes. A cognitive process theory of writing. College Composition and Communication, 32(4):365–387, 1981. ISSN 0010096X. URL http://www.jstor.org/stable/356600.
  9. 9.Gemini. Gemini diffusion, 2025. URL https://deepmind.google/models/gemini-diffusion/.
  10. 10.J. Gottweis, W.-H. Weng, A. Daryin, T. Tu, A. Palepu, P. Sirkovic, A. Myaskovsky, F. Weissenberger, K. Rong, R. Tanno, K. Saab, D. Popovici, J. Blum, F. Zhang, K. Chou, A. Hassidim, B. Gokturk, A. Vahdat, P. Kohli, and V. Natarajan. Towards an ai co-scientist. 02 2025. doi: 10.48550/arXiv.2502.18864.
  11. 11.Grok. Grok, 2025. URL https://grok.com/.
  12. 12.J. Guan, W. Wu, Z. Wen, P. Xu, H. Wang, and M. Huang. Amor: A recipe for building adaptable modular knowledge agents through process feedback. 2024. URL https://arxiv.org/abs/2402.01469.
  13. 13.R. Han, Y. Zhang, P. Qi, Y. Xu, J. Wang, L. Liu, W. Y. Wang, B. Min, and V. Castelli. RAG-QA arena: Evaluating domain robustness for long-form retrieval augmented question answering. In Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 4354–4374, Miami, Florida, USA, Nov. 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.249. URL https://aclanthology.org/2024.emnlp-main.249/.
  14. 14.X. Hu, H. Fu, J. Wang, Y. Wang, Z. Li, R. Xu, Y. Lu, Y. Jin, L. Pan, and Z. Lan. Nova: An iterative planning and search approach to enhance novelty and diversity of llm generated ideas. 2024. URL https://arxiv.org/abs/2410.14255.
  15. 15.Y. Ichihara, Y. Jinnai, T. Morimura, K. Abe, K. Ariu, M. Sakamoto, and E. Uchibe. Evaluation of best-of-n sampling strategies for language model alignment. Transactions on Machine Learning Research, 2025. ISSN 2835-8856. URL https://openreview.net/forum?id=H4S4ETc8c9.
  16. 16.B. Jin, H. Zeng, Z. Yue, D. Wang, H. Zamani, and J. Han. Search-r1: Training llms to reason and leverage search engines with reinforcement learning. arXiv preprint arXiv:2503.09516, 2025.
  17. 17.Kimi-Researcher. Kimi-researcher end-to-end rl training for emerging agentic capabilities, 2025. URL https://moonshotai.github.io/Kimi-Researcher/.
  18. 18.K.-H. Lee, I. Fischer, Y.-H. Wu, S. B. Dave Marwood, D. Schuurmans, and X. Chen. Evolving deeper llm thinking. 2025. URL https://arxiv.org/abs/2501.09891.
  19. 19.D. Li, B. Jiang, L. Huang, A. Beigi, C. Zhao, Z. Tan, A. Bhattacharjee, Y. Jiang, C. Chen, T. Wu, K. Shu, L. Cheng, and H. Liu. From generation to judgment: Opportunities and challenges of llm-as-a-judge. arXiv preprint arXiv: 2411.16594, 2024.
  20. 20.X. Li, G. Dong, J. Jin, Y. Zhang, Y. Zhou, Y. Zhu, P. Zhang, and Z. Dou. Search-o1: Agentic search-enhanced large reasoning models. CoRR, abs/2501.05366, 2025a. doi: 10.48550/ARXIV.2501.05366. URL https://doi.org/10.48550/arXiv.2501.05366.
  21. 21.X. Li, J. Jin, G. Dong, H. Qian, Y. Zhu, Y. Wu, J.-R. Wen, and Z. Dou. Webthinker: Empowering large reasoning models with deep research capability. 2025b. URL https://arxiv.org/abs/2504.21776.
  22. 22.T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, Z. Tu, and S. Shi. Encouraging divergent thinking in large language models through multi-agent debate. arXiv preprint arXiv:2305.19118, 2023.
  23. 23.A. Lim, S. Jain, and V. Seng. Deepconsult: A deep research benchmark for consulting / business queries, 2025. URL https://github.com/Su-Sea/ydc-deep-research-evals.
  24. 24.Y. Liu, H. Zhou, Z. Guo, E. Shareghi, I. Vulic, A. Korhonen, and N. Collier. Aligning with human judgement: The role of pairwise preference in large language model evaluators. arXiv preprint arXiv:2403.16950, 2024.
  25. 25.C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha. The AI Scientist: Towards fully automated open-ended scientific discovery. arXiv preprint arXiv:2408.06292, 2024.
  26. 26.A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark. Self-refine: Iterative refinement with self-feedback. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=S37hOerQLB.
  27. 27.G. Mialon, C. Fourrier, C. Swift, T. Wolf, Y. LeCun, and T. Scialom. Gaia: a benchmark for general ai assistants. 11 2023. URL https://arxiv.org/abs/2311.12983.
  28. 28.S. Nie, F. Z. 1, Z. You, X. Zhang, J. Ou, J. Hu, J. Zhou, Y. Lin, J.-R. Wen, and C. Li. Large language diffusion models. 2025. URL https://arxiv.org/abs/2502.09992.
  29. 29.A. Novikov, N. Vu, M. Eisenberger, E. Dupont, P.-S. Huang, A. Z. Wagner, S. Shirobokov, B. Kozlovskii, F. J. R. Ruiz, M. P. K. Abbas Mehrabian, A. See, S. Chaudhuri, G. Holland, A. Davies, S. Nowozin, P. Kohli, and M. Balog. Alphaevolve: A coding agent for scientific and algorithmic discovery. 2025. URL https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdf.
  30. 30.OpenAI. Introducing deep research, 2025. URL https://openai.com/index/introducing-deep-research/.
  31. 31.Perplexity. Introducing perplexity deep research, 2025. URL https://www.perplexity.ai/hub/blog/introducing-perplexity-deep-research.
  32. 32.L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, M. Choi, A. Agrawal, A. Chopra, A. Khoja, R. Kim, R. Ren, J. Hausenloy, O. Zhang, M. Mazeika, T. Nguyen, D. Anderson, I. A. Shah, M. Doroshenko, A. C. Stokes, M. Mahmood, J. Lee, O. Pokutnyi, O. Iskra, J. P. Wang, R. Gerbicz, J.-C. Levin, S. Popov, F. Feng, S. Y. Feng, H. Zhao, M. Yu, V. Gangal, C. Zou, Z. Wang, M. Kazakov, G. Galgon, J. Schmitt, A. Sanchez, Y. Lee, W. Yeadon, S. Sauers, M. Roth, C. Agu, S. Riis, F. Giska, S. Utpala, A. Cheatom, Z. Giboney, G. M. Goshu, S.-J. Crowson, M. M. Naiya, N. Burns, L. Finke, Z. Cheng, H. Park, F. Fournier-Facio, J. Zampese, J. Wydallis, J. B. Wydallis, R. G. Hoerr, M. Nandor, T. Gehrunger, J. Cai, B. McCarty, J. Nam, E. Taylor, J. Jin, G. A. Loume, H. Cao, A. C. Garretson, D. Sileo, Q. Ren, D. Cojoc, P. Arkhipov, U. Qazi, A. Bacho, A. Li, S. Motwani, C. S. de Witt, A. Kopylov, J. Veith, E. Singer, P. Rissone, J. Jin, J. W. L. Shi, C. G. Willcocks, A. Prabhu, L. Tang, K. Zhou, E. de Oliveira Santos, A. P. Maksimov, E. Vendrow, K. Zenitani, J. Robinson, A. Mikov, J. Guillod, Y. Li, B. Pageler, J. Vendrow, V. Kuchkin, P. Marion, D. Efremov, J. Lynch, K. Liang, A. Gritsevskiy, D. Martinez, N. Crispino, D. Zvonkine, N. W. Fraga, S. Soori, O. Press, H. Tang, J. Salazar, S. R. Green, L. Brüssel, M. Twayana, A. Dieuleveut, T. R. Rogers, W. Zhang, R. Finocchio, B. Li, J. Yang, A. Rao, G. Loiseau, M. Kalinin, M. Lukas, C. Manolescu, N. Stambaugh, S. Mishra, A. G. K. Kamdoum, T. Hogg, A. Jin, C. Bosio, G. Sun, B. P. Coppola, H. Heidinger, R. Sayous, S. Ivanov, J. M. Cavanagh, J. Shen, J. M. Imperial, P. Schwaller, S. Senthilkuma, A. M. Bran, A. Algaba, B. Verbeken, K. V. den Houte, L. V. D. Sypt, D. Noever, L. Schut, I. Sucholutsky, E. Zheltonozhskii, Q. Yuan, D. Lim, R. Stanley, S. Sivarajan, T. Yang, J. Maar, J. Wykowski, M. Oller, J. Sandlin, A. Sahu, C. G. Ardito, Y. Hu, F. M. Dias, T. Kreiman, K. Rawal, T. G. Vilchis, Y. Zu, M. Lackner, J. Koppel, J. Nguyen, D. S. Antonenko, S. Chern, B. Zhao, P. Arsene, S. Ivanov, R. Poświata, C. Wang, D. Li, D. Crisostomi, A. Dehghan, A. Achilleos, J. A. Ambay, B. Myklebust, A. Sen, D. Perrella, N. Kaparov, M. H. Inlow, A. Zang, K. Ramakrishnan, D. Orel, V. Poritski, S. Ben-David, Z. Berger, P. Whitfill, M. Foster, D. Munro, L. Ho, D. B. Hava, A. Kuchkin, R. Lauff, D. Holmes, F. Sommerhage, A. Zhang, R. Moat, K. Schneider, D. Pyda, Z. Kazibwe, M. Singh, D. Clarke, D. H. Kim, S. Fish, V. Elser, V. E. G. Vilchis, I. Klose, C. Demian, U. Anantheswaran, A. Zweiger, G. Albani, J. Li, N. Daans, M. Radionov, V. Rozhoň, V. Ginis, Z. Ma, C. Stump, J. Platnick, V. Nevirkovets, L. Basler, M. Piccardo, N. Cohen, V. Singh, J. Tkadlec, P. Rosu, A. Goldfarb, P. Padlewski, S. Barzowski, K. Montgomery, A. Menezes, A. Patel, Z. Wang, J. Tucker-Foltz, J. Stade, D. Grabb, T. Goertzen, F. Kazemi, J. Milbauer, A. Shukla, H. Elgnainy, Y. C. L. Labrador, H. He, L. Zhang, A. Givré, H. Wolff, G. Demir, M. F. Aziz, Y. Kaddar, I. Ängquist, Y. Chen, E. Thornley, R. Zhang, J. Pan, A. Terpin, N. Muennighoff, H. Schoelkopf, E. Zheng, A. Carmi, J. Shah, E. D. L. Brown, K. Zhu, M. Bartolo, R. Wheeler, A. Ho, S. Barkan, J. Wang, M. Stehberger, E. Kretov, P. Bradshaw, J. Heimonen, K. Sridhar, Z. Hossain, I. Akov, Y. Makarychev, J. Tam, H. Hoang, D. M. Cunningham, V. Goryachev, D. Patramanis, M. Krause, A. Redenti, D. Aldous, J. Lai, S. Coleman, J. Xu, S. Lee, I. Magoulas, S. Zhao, N. Tang, M. K. Cohen, M. Carroll, O. Paradise, J. H. Kirchner, S. Steinerberger, M. Ovchynnikov, J. O. Matos, A. Shenoy, M. Wang, Y. Nie, P. Giordano, P. Petersen, A. Sztyber-Betley, P. Faraboschi, R. Riblet, J. Crozier, S. Halasyamani, A. Pinto, S. Verma, P. Joshi, E. Meril, Z.-X. Yong, A. Tee, J. Andréoletti, O. Weller, R. Singhal, G. Zhang, A. Ivanov, S. Khoury, N. Gustafsson, H. Mostaghimi, K. Thaman, Q. Chen, T. Q. Khánh, J. Loader, S. Cavalleri, H. Szlyk, Z. Brown, H. Narayan, J. Roberts, W. Alley, K. Sun, R. Stendall, M. Lamparth, A. Reuel, T. Wang, H. Xu, P. Hernández-Cámara, F. Martin, T. Preu, T. Korbak, M. Abramovitch, D. Williamson, I. Bosio, Z. Chen, B. Bálint, E. J. Y. Lo, M. I. S. Nunes, Y. Jiang, M. S. Bari, P. Kassani, Z. Wang, B. Ansarinejad, Y. Sun, S. Durand, G. Douville, D. Tordera, G. Balabanian, E. Anderson, L. Kvistad, A. J. Moyano, H. Milliron, A. Sakor, M. Eron, I. C. McAlister, A. F. D. O., S. Shah, X. Zhou, F. Kamalov, R. Clark, S. Abdoli, T. Santens, H. K. Wang, E. Chen, A. Tomasiello, G. B. D. Luca, S.-Z. Looi, V.-K. Le, N. Kolt, N. Mündler, A. Semler, E. Rodman, J. Drori, C. J. Fossum, L. Gloor, M. Jagota, R. Pradeep, H. Fan, T. Shah, J. Eicher, M. Chen, K. Thaman, W. Merrill, M. Firsching, C. Harris, S. Ciobâcă, J. Gross, R. Pandey, I. Gusev, A. Jones, S. Agnihotri, P. Zhelnov, S. Usawasutsakorn, M. Mofayezi, A. Piperski, M. Carauleanu, D. K. Zhang, K. Dobarskyi, D. Ler, R. Leventov, I. Soroko, T. Jansen, S. Creighton, P. Lauer, J. Duersch, V. Taamazyan, D. Bezzi, W. Morak, W. Ma, W. Held, T. Ðuc Huy, R. Xian, A. R. Zebaze, M. Mohamed, J. N. Leser, M. X. Yuan, L. Yacar, J. Lengler, K. Olszewska, H. Shahrtash, E. Oliveira, J. W. Jackson, D. E. Gonzalez, A. Zou, M. Chidambaram, T. Manik, H. Haffenden, D. Stander, A. Dasouqi, A. Shen, E. Duc, B. Golshani, D. Stap, M. Uzhou, A. B. Zhidkovskaya, L. Lewark, M. O. Rodriguez, M. Vincze, D. Wehr, C. Tang, S. Phillips, F. Samuele, J. Muzhen, F. Ekström, A. Hammon, O. Patel, F. Farhidi, G. Medley, F. Mohammadzadeh, M. Peñaflor, H. Kassahun, A. Friedrich, C. Sparrow, R. H. Perez, T. Sakal, O. Dhamane, A. K. Mirabadi, E. Hallman, K. Okutsu, M. Battaglia, M. Maghsoudimehrabani, A. Amit, D. Hulbert, R. Pereira, S. Weber, Handoko, A. Peristyy, S. Malina, S. Albanie, W. Cai, M. Mehkary, R. Aly, F. Reidegeld, A.-K. Dick, C. Friday, J. Sidhu, H. Shapourian, W. Kim, M. Costa, H. Gurdogan, B. Weber, H. Kumar, T. Jiang, A. Agarwal, C. Ceconello, W. S. Vaz, C. Zhuang, H. Park, A. R. Tawfeek, D. Aggarwal, M. Kirchhof, L. Dai, E. Kim, J. Ferret, Y. Wang, M. Yan, K. Burdzy, L. Zhang, A. Franca, D. T. Pham, K. Y. Loh, J. Robinson, A. Jackson, S. Gul, G. Chhablani, Z. Du, A. Cosma, J. Colino, C. White, J. Votava, V. Vinnikov, E. Delaney, P. Spelda, V. Stritecky, S. M. Shahid, J.-C. Mourrat, L. Vetoshkin, K. Sponselee, R. Bacho, F. de la Rosa, X. Li, G. Malod, L. Lang, J. Laurendeau, D. Kazakov, F. Adesanya, J. Portier, L. Hollom, V. Souza, Y. A. Zhou, J. Degorre, Y. Yalın, G. D. Obikoya, L. Arnaboldi, Rai, F. Bigi, M. C. Boscá, O. Shumar, K. Bacho, P. Clavier, G. Recchia, M. Popescu, N. Shulga, N. M. Tanwie, D. Peskoff, T. C. H. Lux, B. Rank, C. Ni, M. Brooks, A. Yakimchyk, Huanxu, Liu, O. Häggström, E. Verkama, H. Gundlach, L. Brito-Santana, B. Amaro, V. Vajipey, R. Grover, Y. Fan, G. P. R. e Silva, L. Xin, Y. Kratish, J. Łucki, W.-D. Li, S. Gopi, A. Caciolai, J. Xu, K. J. Scaria, F. Vargus, F. Habibi, Long, Lian, E. Rodolà, J. Robins, V. Cheng, T. Fruhauff, B. Raynor, H. Qi, X. Jiang, B. Segev, J. Fan, S. Martinson, E. Y. Wang, K. Hausknecht, M. P. Brenner, M. Mao, X. Zhang, D. Avagian, E. J. Scipio, A. Ragoler, J. Tan, B. Sims, R. Plecnik, A. Kirtland, O. F. Bodur, D. P. Shinde, Z. Adoul, M. Zekry, A. Karakoc, T. C. B. Santos, S. Shamseldeen, L. Karim, A. Liakhovitskaia, N. Resman, N. Farina, J. C. Gonzalez, G. Maayan, S. Hoback, R. D. O. Pena, G. Sherman, E. Kelley, H. Mariji, R. Pouriamanesh, W. Wu, S. Mendoza, I. Alarab, J. Cole, D. Ferreira, B. Johnson, M. Safdari, L. Dai, S. Arthornthurasuk, A. Pronin, J. Fan, A. Ramirez-Trinidad, A. Cartwright, D. Pottmaier, O. Taheri, D. Outevsky, S. Stepanic, S. Perry, L. Askew, R. A. H. Rodríguez, A. M. R. Minissi, S. Ali, R. Lorena, K. Iyer, A. A. Fasiludeen, S. M. Salauddin, M. Islam, J. Gonzalez, J. Ducey, M. Somrak, V. Mavroudis, E. Vergo, J. Qin, B. Borbás, E. Chu, J. Lindsey, A. Radhakrishnan, A. Jallon, I. M. J. McInnis, P. Kumar, L. P. Goswami, D. Bugas, N. Heydari, F. Jeanplong, A. Apronti, A. Galal, N. Ze-An, A. Singh, J. of Arc Xavier, K. P. Agarwal, M. Berkani, B. A. de Oliveira Junior, D. Malishev, N. Remy, T. D. Hartman, T. Tarver, S. Mensah, J. Gimenez, R. G. Montecillo, R. Campbell, A. Sharma, K. Meer, X. Alapont, D. Patil, R. Maheshwari, A. Dendane, P. Shukla, S. Bogdanov, S. Möller, M. R. Siddiqi, P. Saxena, H. Gupta, I. Enyekwe, R. P. V, Z. EL-Wasif, A. Maksapetyan, V. Rossbach, C. Harjadi, M. Bahaloohoreh, S. Bian, J. Lai, J. L. Uro, G. Bateman, M. Sayed, A. Menshawy, D. Duclosel, Y. Jain, A. Aaron, M. Tiryakioglu, S. Siddh, K. Krenek, A. Hoover, J. McGowan, T. Patwardhan, S. Yue, A. Wang, and D. Hendrycks. Humanity’s last exam, 2025. URL https://arxiv.org/abs/2501.14249.
  33. 33.J. Qiu, X. Qi, T. Zhang, X. Juan, J. Guo, Y. Lu, Y. Wang, Z. Yao, Q. Ren, X. Jiang, X. Zhou, D. Liu, L. Yang, Y. Wu, K. Huang, S. Liu, H. Wang, and M. Wang. Alita: Generalist agent enabling scalable agentic reasoning with minimal predefinition and maximal self-evolution. arXiv preprint arXiv:2505.20286, 2025.
  34. 34.O. D. Research. Open deep research, 2025. URL https://github.com/langchain-ai/open_deep_research.
  35. 35.G. Researcher. Gpt researcher, 2025. URL https://github.com/assafelovic/gpt-researcher.
  36. 36.A. Roucher, A. V. del Moral, merve, T. Wolf, and C. Fourrier. Open-source deepresearch – freeing our search agents, 2025. URL https://huggingface.co/blog/open-deep-research.
  37. 37.S. Schmidgall, Y. Su, Z. Wang, X. Sun, J. Wu, X. Yu, J. Liu, Z. Liu, and E. Barsoum. Agent laboratory: Using llm agents as research assistants. 2025. URL https://arxiv.org/abs/2501.04227.
  38. 38.H. Shen, J. Zhang, B. Xiong, R. Hu, S. Chen, Z. Wan, X. Wang, Y. Zhang, Z. Gong, G. Bao, et al. Efficient diffusion models: A survey. Transactions on Machine Learning Research (TMLR), 2025.
  39. 39.W. Shi, H. Tan, C. Kuang, X. Li, X. Ren, C. Zhang, H. Chen, Y. Wang, L. Shang, F. Yu, and Y. Wang. Pangu deepdiver: Adaptive search intensity scaling via open-web reinforcement learning. arXiv preprint arXiv:2505.24332, 2025.
  40. 40.C. Si, D. Yang, and T. Hashimoto. Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers. 2024. URL https://arxiv.org/abs/2409.04109.
  41. 41.I. Stelmakh, Y. Luan, B. Dhingra, and M.-W. Chang. ASQA: Factoid questions meet long-form answers. In Y. Goldberg, Z. Kozareva, and Y. Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8273–8288, Abu Dhabi, United Arab Emirates, Dec. 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.emnlp-main.566. URL https://aclanthology.org/2022.emnlp-main.566/.
  42. 42.J. Tang, L. Xia, Z. Li, and C. Huang. Ai-researcher: Autonomous scientific innovation. 2025. URL https://arxiv.org/abs/2505.18705.
  43. 43.H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal. Musique: Multihop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10:539–554, 2022. doi: 10.1162/tacl_a_00475. URL https://aclanthology.org/2022.tacl-1.31/.
  44. 44.J. Wei, X. Wang, D. Schuurmans, M. Bosma, b. ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou. Chain-of-thought prompting elicits reasoning in large language models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc., 2022. URL https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf.
  45. 45.Y. Yamada, R. T. Lange, C. Lu, S. Hu, C. Lu, J. Foerster, J. Clune, and D. Ha. The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search. 2025. URL https://arxiv.org/abs/2504.08066.
  46. 46.L. Yang, Z. Zhang, Y. Song, S. Hong, R. Xu, Y. Zhao, W. Zhang, B. Cui, and M.-H. Yang. Diffusion models: A comprehensive survey of methods and applications. 2022. URL https://arxiv.org/abs/2209.00796.
  47. 47.Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii, editors, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium, Oct.-Nov. 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1259. URL https://aclanthology.org/D18-1259/.
  48. 48.J. Yoon, H. Cho, Y. Bengio, and S. Ahn. Fast monte carlo tree diffusion: 100x speedup via parallel sparse planning. 06 2025. URL https://arxiv.org/abs/2506.09498.
  49. 49.K. Zhang, X. Yang, W. Y. Wang, and L. Li. Redi: efficient learning-free diffusion inference via trajectory retrieval. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org, 2023.
  50. 50.Y. Zheng, S. Sun, L. Qiu, D. Ru, C. Jiayang, X. Li, J. Lin, B. Wang, Y. Luo, R. Pan, Y. Xu, Q. Min, Z. Zhang, Y. Wang, W. Li, and P. Liu. OpenResearcher: Unleashing AI for accelerated scientific research. In D. I. Hernandez Farias, T. Hope, and M. Li, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 209–218, Miami, Florida, USA, Nov. 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-demo.22. URL https://aclanthology.org/2024.emnlp-demo.22/.
  51. 51.Y. Zheng, D. Fu, X. Hu, X. Cai, L. Ye, P. Lu, and P. Liu. Deepresearcher: Scaling deep research via reinforcement learning in real-world environments. 2025. URL https://arxiv.org/abs/2504.03160.
  52. 52.M. Świechowski, K. Godlewski, B. Sawicki, and J. Mańdziuk. Monte carlo tree search: a review of recent modifications and applications. Artificial Intelligence Review, 56, 07 2022. doi: 10.1007/s10462-022-10228-y.

Citation

MLA
Han, R., et al. “Deep Researcher with Test-Time Diffusion”. arXiv, 2025, http://arxiv.org/abs/2507.16075v1.
APA
Han, R., Chen, Y., CuiZhu, Z., Miculicich, L., Sun, G., Bi, Y., Wen, W., Wan, H., Wen, C., Maître, S., Lee, G., Tirumalashetty, V., Xue, E., Zhang, Z., Haykal, S., Gokturk, B., Pfister, T., & Lee, C.-Y. (2025). Deep Researcher with Test-Time Diffusion. arXiv. http://arxiv.org/abs/2507.16075v1
Chicago
Han, R., Y. Chen, Z. CuiZhu, et al. 2025. “Deep Researcher with Test-Time Diffusion”. arXiv. http://arxiv.org/abs/2507.16075v1.
Harvard
Han, R. et al. (2025) “Deep Researcher with Test-Time Diffusion”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2507.16075v1.
Vancouver
1. Han R, Chen Y, CuiZhu Z, et al (2025) Deep Researcher with Test-Time Diffusion. arXiv

BibTeX

@article{han2025deep,
  title = {Deep Researcher with Test-Time Diffusion},
  author = {Han, Rujun and Chen, Yanfei and CuiZhu, Zoey and Miculicich, Lesly and Sun, Guan and Bi, Yuanjun and Wen, Weiming and Wan, Hui and Wen, Chunfeng and Maître, Solène and Lee, George and Tirumalashetty, Vishy and Xue, Emily and Zhang, Zizhao and Haykal, Salem and Gokturk, Burak and Pfister, Tomas and Lee, Chen-Yu},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2507.16075v1},
  eprint = {2507.16075}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/