Deep Researcher with Test-time Diffusion
Rujun HanYanfei ChenGuan SunLesly MiculicichZoey CuiZhuYuanjun (Sophia) BiWeiming WenHui WanChunfeng WenSolène Maître
Proposes Test-Time Diffusion Deep Researcher, a framework that models long-form report generation as an iterative diffusion process to refine structured drafts through continuous retrieval, outperforming existing LLM research agents on complex multi-hop reasoning benchmarks.
Autonomous research systems powered by large language models have advanced rapidly, yet their ability to generate complex, long-form research reports frequently hits a ceiling. Most existing systems operate through rigid, linear, or disconnected search workflows that struggle to maintain broad context, resulting in significant information loss and uncoordinated findings. The article introduces and evaluates Test-Time Diffusion Deep Researcher (TTD-DR), a framework designed to mimic iterative human writing behavior by framing report generation as a diffusion process that continuously refines an initial draft with external web search and component-level optimization.
The framework operates across three key stages: generating a research plan, conducting iterative information retrieval, and synthesizing a final report. Credibility and robustness are established through two core mechanisms. First, report-level denoising uses the evolving draft to dynamically guide search queries while immediately integrating retrieved findings back into the text across up to 20 revision steps. Second, component-wise self-evolution optimizes individual sub-tasks—such as planning, query generation, and answer synthesis—by generating diverse variants, scoring them with an automated evaluator, refining them based on critique, and merging the best elements. The system was evaluated across long-form report benchmarks (LongForm Research and DeepConsult) and complex multi-hop reasoning datasets (Humanity's Last Exam text subsets and GAIA), benchmarked against proprietary and open-source systems including OpenAI Deep Research, Perplexity Deep Research, and Grok DeeperSearch.
The evaluation demonstrates several key findings:
- TTD-DR achieves state-of-the-art performance in long-form report generation, recording a 69.1% win rate against OpenAI Deep Research on the LongForm Research benchmark and a 74.5% win rate on DeepConsult.
- In multi-hop reasoning and concise question answering, TTD-DR outperforms OpenAI Deep Research by 4.8 percentage points on search-intensive academic queries (33.9% vs. 29.1%) and by 7.7 percentage points across the broader academic benchmark (34.3% vs. 26.6%), while also leading on real-world general tasks (69.1% vs. 67.4%).
- Denoising with retrieval increases search query novelty by more than 12 percentage points and captures over 51% of the final report's key information by the ninth search step, outperforming 20 steps of self-evolution alone.
- Latency-efficiency analyses show that TTD-DR exhibits a steeper efficiency frontier than baseline methods, yielding greater output quality per unit of processing time.
These findings indicate that adopting human-like drafting and continuous revision loops fundamentally improves coherence, depth, and factual retention in automated research. Unlike conventional systems that perform isolated searches before writing, maintaining a central working draft reduces information loss and prevents circular or redundant queries. For organizations seeking automated strategic analysis, market intelligence, or technical synthesis, this architecture provides a viable pathway to higher-quality outputs using standard search engines without relying on opaque, proprietary toolstacks.
Organizations developing or deploying automated research agents should consider transitioning from linear multi-agent pipelines to draft-centric, iterative revision frameworks. Engineering teams should prioritize test-time scaling methods that combine sub-task self-evolution with dynamic retrieval loops. However, decision-makers should note that this study focused solely on text-based web search; the current framework does not integrate code execution or multimodal web browsing environments. Additionally, because evaluations rely on automated model judges calibrated to human preferences, critical strategic deployments should maintain human-in-the-loop oversight until further real-world pilot validations are completed.
- Paper: Diffusion-LM Improves Controllable Text Generation, Xiang Lisa Li et al. (2022). Diffusion-LM establishes how iterative denoising can generate text, the key generative premise behind TTD-DR’s diffusion framing.
- Paper: DiffusER: Discrete Diffusion via Edit-based Reconstruction, Machel Reid et al. (2023). DIFFUSER develops text generation as repeated edits to an existing draft, clarifying the revision-based process that TTD-DR adapts for research reports.
- Paper: Active Retrieval Augmented Generation, Zhengbao Jiang et al. (2023). FLARE shows how retrieval can be triggered and refreshed during long-form generation, providing a direct foundation for TTD-DR’s retrieval-informed refinement.
- Paper: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection, Akari Asai et al. (2024). Self-RAG links retrieval with critique and generation, useful groundwork for understanding TTD-DR’s retrieval-informed agent workflow.
- Paper: General Agentic Memory Via Deep Research, B. Y. Yan et al. (2025). General Agentic Memory extends iterative research-agent workflows to long-term context construction, making it a natural next step after TTD-DR’s retrieval-driven report refinement.
- Paper: A Benchmark for Deep Information Synthesis, Debjit Paul et al. (2026). A Benchmark for Deep Information Synthesis carries deep-research evaluation toward fragmented real-world sources and multi-step analytical synthesis.
- Paper: PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing, Yiwen Song et al. (2026). PaperOrchestra applies coordinated search, outlining, drafting, and iterative review to automate complete research papers, extending TTD-DR’s report-writing approach.
