AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery

Guiyao TieJiawen ShiDingjie SongYixiao HuangZiji ShengXueyang ZhouDaizong LiuPan ZhouYongchao ChenRan Xu

article2026arXiv0 citations

Establishes a comprehensive framework for AI-powered scientific discovery by analyzing automated research workflows across five stages and defining concrete evaluation criteria for novelty, validity, reliability, and provenance.

Listen

Artificial intelligence is undergoing a major transition from isolated, task-specific scientific tools—such as molecular property predictors and narrow retrieval systems—toward workflow-level scientific research automation. While existing tools can perform bounded tasks, generate plausible ideas, and produce polished papers, the field remains fragmented across different levels of autonomy and execution environments. To address these emerging developments, the article establishes a unified analytical framework called AutoResearch. The article's main objective is to evaluate the developmental spectrum of AI-driven scientific workflow automation and demonstrate how control, evidence, validation, and accountability are redistributed across the entire discovery cycle.

To conduct this evaluation, the article synthesizes historical and contemporary research systems, mixed-initiative co-research frameworks, domain-specific implementations, and benchmarking ecosystems across five core workflow stages: literature grounding, hypothesis formation and planning, experimentation and tool use, validation and review, and reporting. It defines a five-level autonomy spectrum ranging from fully manual work (Level 0) to human-led assistance (Level 1), human-verified execution (Level 2), AI-led coordination (Level 3), and fully autonomous scientific discovery (Level 4), mapping the practical boundaries and ceilings across diverse scientific disciplines.

The analysis yields four critical findings. First, the article finds that current systems are heavily concentrated in Levels 1 and 2, which represent human-steered and human-verified execution (termed "Vibe Research"); no mature system has achieved reliable Level 3 coordination or Level 4 end-to-end scientific closure. Second, pipeline breadth must not be confused with achieved scientific autonomy: while systems can draft papers and run code, they remain weak at evidence preservation, exception handling, baseline comparison, and rejecting flawed hypotheses. Third, evaluation criteria must shift from task completion alone to five workflow-level credibility dimensions: novelty, validity, impact, reliability, and provenance. Fourth, the ceiling of scientific autonomy is strongly domain-conditioned: computational and formal sciences achieve higher automation because their digital artifacts are executable, replayable, and rapidly verifiable, whereas wet-lab biology, medicine, and social sciences remain constrained by physical embodiment, delayed validation, safety risks, and regulatory accountability.

These findings indicate that deploying automated scientific pipelines without robust validation mechanisms introduces significant risks of compounding errors, ungrounded claims, and irreproducible results. Leaders and stakeholders must understand that high-level paper generation and code execution do not equate to credible scientific discovery. Consequently, the article advises treating current automated systems as advanced productivity aids rather than autonomous research agents. Decision-makers should establish explicit human-in-the-loop validation checkpoints, invest in auditable provenance tracking, and avoid relying on automated acceptance for high-stakes, embodied, or clinical research.

Finally, the conclusions are framed by clear limitations: contemporary automated systems lack standardized cross-domain benchmarks, and reliable evaluation frameworks that simultaneously measure novelty, empirical validity, and operational provenance are still emerging. Readers should maintain high confidence in the potential of AI to accelerate computational workflows, but exercise caution against premature claims of autonomous scientific discovery in complex, empirical, and safety-critical domains.

arXiv: 2605.23204
Cover for AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery

Abstract

Scientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding, hypothesis generation, experimentation, validation, reporting, and revision. This shift marks a transition from task-level AI for science to workflow-level research automation. Yet current systems remain fragmented, differing in autonomy, domain scope, execution environment, validation mechanism, and human oversight, while still struggling with evidence preservation, reproducibility, weak-direction rejection, provenance tracking, cross-domain robustness, and accountable scientific closure. This survey examines these developments through AutoResearch, defined as the developmental spectrum of AI-powered scientific workflow automation. Within it, Vibe Research denotes the human-steered region of prompt-based assistance and human-verified execution, whereas emerging AI-led systems coordinate larger portions of the discovery loop without achieving robust autonomy. We analyze how research systems redistribute control, evidence, execution, validation, and accountability across workflows and organize the field around five workflow conditions: literature and research grounding; hypothesis formation and planning; experimentation and tool use; feedback, validation, and review; and reporting and knowledge communication. We further synthesize AI scientist systems, mixed-initiative co-research frameworks, benchmarks, domain deployments, and open-source infrastructures. Finally, we propose five evaluation dimensions--novelty, validity, impact, reliability, and provenance--and show that AutoResearch autonomy is domain-conditioned, being more credible in structured, executable, and rapidly verifiable settings but limited in embodied, delayed, heterogeneous, ethical, or institutionally accountable contexts.

Table of Contents

  • 1 Introduction
  • 2 Overview of AutoResearch
  • 2.1 History of AutoResearch
  • 2.2 Contemporary Landscape of AutoResearch
  • 3 Technical Foundations of AutoResearch
  • 3.1 Workflow Conditions for Scientific Autonomy
  • 3.2 Stage I: Literature and Research Grounding
  • 3.3 Stage II: Hypothesis Formation and Planning
  • 3.4 Stage III: Experimentation and Tool Use
  • 3.5 Stage IV: Feedback, Validation, and Review
  • 3.6 Stage V: Reporting and Knowledge Communication
  • 4 Evaluation of AutoResearch
  • 4.1 Evaluative Burdens Across the AutoResearch Spectrum
  • 4.2 Scientific Quality and Autonomy Assessment
  • 4.3 Evaluation Instruments and Benchmark Landscape
  • 5 Domains of AutoResearch
  • 5.1 Computational and Formal Sciences
  • 5.2 Physical Sciences and Engineering
  • 5.3 Embodied Intelligence
  • 5.4 Chemistry and Materials
  • 5.5 Biology and Biomedicine
  • 5.6 Medicine and Clinical Research
  • 5.7 Economics and Social Sciences
  • 5.8 Earth and Environmental Sciences
  • 6 Discussion
  • 6.1 Rethinking the Capabilities of AutoResearch
  • 6.2 Evaluation for AutoResearch
  • 6.3 The Generalization Gap: Beyond Computational and Formal Sciences
  • 6.4 Reliability, Trustworthiness, and Auditability of AutoResearch
  • 6.5 Ethical and Societal Implications
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Five-Level Scientific Workflow Autonomy Spectrum (L0–L4)

    definition

    The scientific workflow autonomy of AI systems in research automation is categorized across five hierarchical levels (L0L0 to L4L4), defined by the allocation of workflow control, task execution, validation authority, and scientific accountability:

    • L0L0: Human Only: Scientific inquiry is entirely human-led, human-executed, and human-verified. Human researchers formulate problems, design and run experiments, evaluate evidence, and retain full accountability at every step. Digital tools provide local assistance without shifting scientific agency.
    • L1L1: Human-Led, AI-Assisted: The workflow remains human-led, with AI providing bounded cognitive assistance (e.g., literature search, summarization, drafting, exploratory ideation). Humans decide research directions, execute tasks, validate findings, and retain full control.
    • L2L2: Human-Verified, AI-Executed: AI executes substantive multi-step research operations (e.g., file modification, code synthesis, tool execution, intermediate data analysis, running research pipelines). However, final acceptance, scientific validity assessment, branch filtering, and accountability remain human-held.
    • L3L3: AI-Led, Human-Assisted: AI organizes and coordinates major spans of the research loop (literature grounding, planning, execution, validation, revision, and reporting). Humans shift from routine stepwise verification to high-level supervision, exception handling, and edge-case interventions when the automated loop encounters uncertainty.
    • L4L4: AI-Autonomous: AI carries out end-to-end scientific discovery workflows autonomously, achieving routine scientific closure (problem formulation, hypothesis grounding, execution, validation, rejection of unpromising directions, and provenance preservation) without structural reliance on human execution, subject to post hoc institutional governance.

    The framework designates the L1L1–L2L2 human-steered region as Vibe Research, distinguishing it from the stricter, future-facing L3L3–L4L4 AutoResearch autonomy frontier.

  2. Knowl 2 — Sub-Categorization of Level 2 (L2) Human-Verified AI Execution

    definition

    Within the AutoResearch framework, Level 2 (L2L2: Human-Verified, AI-Executed) is refined into three distinct operational regimes based on workflow span and interaction structure:

    • L2-SL2\text{-S} (Single-Step Automated Execution): AI carries out bounded, well-specified discrete operations, such as executing specific tool calls, running local scripts, operating laboratory hardware protocols, training individual models, or performing isolated data analyses under strict human direction and immediate verification.
    • L2-IL2\text{-I} (Interactive Workflow Automation): AI coordinates multi-step scientific workflows via mixed-initiative interaction, incorporating collaborative ideation, step-by-step steering, iterative feedback, and human checkpoints across multiple actions.
    • L2-PL2\text{-P} (Pipeline Automation under Human Verification): AI connects consecutive scientific stages—spanning literature grounding, hypothesis generation, code implementation, experiment execution, result analysis, and paper drafting—into an integrated loop. Despite broad pipeline coverage, the system remains classified as L2L2 because human verification remains structurally essential to judge scientific validity, novelty, reproducibility, and real-world acceptance.
  3. Knowl 3 — Five-Stage Workflow Architecture of AutoResearch

    model/method

    AutoResearch systems decompose scientific inquiry into five recurring technical workflow stages, each imposing specific constraints on the discovery process:

    1. Stage I: Literature and Research Grounding: Accesses, filters, extracts, and organizes prior scientific work to transform raw literature into structured, traceable evidence states that constrain downstream reasoning.
    2. Stage II: Hypothesis Formation and Planning: Converts grounded context into candidate research directions, operationalizing problem formulation, task decomposition, branch expansion, feasibility analysis, and candidate prioritization before resource commitment.
    3. Stage III: Experimentation and Tool Use: Translates candidate plans into actionable interventions by coupling them to computational runtimes, external tools, simulators, robotic laboratory instruments, or human checkpoints to produce empirical resistance and execution artifacts.
    4. Stage IV: Feedback, Validation, and Review: Subjects intermediate and final execution outputs to rejection pressure via consistency reruns, baseline comparisons, error detection, critic panels, or expert evaluations to filter out weak or invalid claims.
    5. Stage V: Reporting and Knowledge Communication: Translates workflow states and verified findings into communicable scientific artifacts (manuscripts, visualizations, tables, code repositories, metadata, and provenance links) while maintaining strict alignment between claims and underlying evidence.
  4. Knowl 4 — Literature and Research Grounding Regimes

    model/method

    Literature and research grounding (Stage I of AutoResearch) operates across four distinct technical regimes categorized by evidential persistence and structural depth:

    1. Search-Centered Grounding: Employs query reformulation, keyword/semantic retrieval, reranking, and summarization to construct local textual context for immediate drafting or framing. Evidential persistence is low and context remains stage-local.
    2. Evidence-Centered Grounding: Inserts an explicit claim-evidence binding layer between retrieval and downstream generation. It extracts specific passages and builds citation-backed rationales and evidence ledgers to substantiate individual claims.
    3. Structure-Centered Grounding: Extracts scientific entities, concepts, and relations (e.g., method, dataset, mechanism, result, limitation) from literature corpora into ontological knowledge graphs, supporting relation-aware gap detection and interdisciplinary analogy reasoning.
    4. Literature-Memory Grounding: Preserves screened papers as indexed metadata, evidence cards (specifying claims, methods, results, and limitations), citation graphs, and prior-art boundary notes. This creates a persistent, reusable evidence memory queried across all subsequent research stages.
  5. Knowl 5 — Hypothesis Formation and Planning Regimes

    model/method

    Hypothesis formation and planning (Stage II of AutoResearch) converts grounded evidence into candidate research directions through four technical regimes:

    1. Proposal-Centered Ideation: A centralized controller generates an initial candidate proposal from grounded context, performs local refinement, and outputs a ranked hypothesis or plan sketch.
    2. Deliberative Multi-Agent Ideation: Distributes ideation across specialized agents (e.g., generation, critique, ranking, evolution, meta-review agents) that engage in multi-round debate, comparative critique, and iterative candidate refinement before execution.
    3. Structure-Guided Ideation: Constrains hypothesis formation using structured prior knowledge (e.g., knowledge graphs or research-paper facets such as purpose, mechanism, and evaluation), formulating directions through explicit structural recombination and gap analysis.
    4. Search-Based and Evolutionary Planning: Treats planning as tree search or evolutionary optimization over a candidate space. Hypotheses are branched, scored by evaluators (against novelty, feasibility, expected gain, and resource/safety constraints), pruned, and iteratively recombined.
  6. Knowl 6 — Experimentation and Tool Use Execution Regimes

    model/method

    Experimentation and tool use (Stage III of AutoResearch) translates research plans into actionable procedures through four operational regimes:

    1. Code-Native Execution: Operates inside software repositories via code-editing loops (inspecting files, localizing bugs, applying patches, and executing runtime/unit tests) to produce executable scripts, run logs, and patch packages.
    2. Tool-Orchestrated Execution: Decomposes plans into parameter-filled calls to external scientific software, domain databases, computational calculators, simulators, and domain APIs, assembling observations into downstream workflows.
    3. Laboratory-Robotic Execution: Compiles experimental recipes into machine protocols, schedules robotic hardware, conducts physical experiments/measurements, and collects characterization data within self-driving laboratories.
    4. Human-Gated Execution: Imposes explicit human-in-the-loop checkpoints before consequential actions. Proposed operations pass through validity, safety, resource, and ethics gates where human experts approve, revise, or block execution.
  7. Knowl 7 — Feedback, Validation, and Rejection Regimes

    model/method

    Feedback, validation, and review (Stage IV of AutoResearch) introduces discriminative rejection pressure through three mechanisms:

    1. Execution-Coupled Rerun Validation: Applies immediate empirical resistance by rerunning code, performing ablation studies, conducting baseline comparisons, and executing reproducibility scripts to filter out execution bugs and noisy metrics.
    2. Critique-Mediated Validation: Employs automated reviewer agents or multi-agent critique panels that evaluate candidate results against rubrics for methodological soundness, novelty, clarity, and missing controls, producing structured revision signals.
    3. Expert- or Temporally-Grounded Validation: Evaluates findings through human domain-expert review, delayed longitudinal follow-up, or full-lifecycle scientific rediscovery benchmarks, testing whether candidate claims withstand real-world scientific standards beyond one-shot execution success.
  8. Knowl 8 — Reporting and Knowledge Communication Regimes

    model/method

    Reporting and knowledge communication (Stage V of AutoResearch) structures scientific dissemination through three distinct communication regimes:

    1. Draft-Centered Reporting: Transforms workflow state and evidence context into fluent long-form manuscripts and section drafts (e.g., related work, methodology, discussion) via sequential language generation and local summarization.
    2. Review-Centered Communication: Implements a dialogic reporting process where candidate manuscripts undergo automated peer review, followed by author-agent response generation, counter-argumentation, and targeted manuscript revision.
    3. Artifact-Linked Reporting: Packages the final manuscript in tight alignment with inspectable and executable artifacts, ensuring explicit bidirectional synchronization between textual claims, tables, figures, source code, datasets, and metadata provenance.
  9. Knowl 9 — Dual-Target AutoResearch Evaluation Framework

    model/method

    A complete evaluation of AutoResearch requires evaluating both Scientific Quality (epistemic credibility of outputs) and Autonomy Profile (degree of independent system agency):

    1. Scientific Quality Dimensions

    • Novelty: Non-obviousness relative to prior scientific literature, evaluated temporally and against expert-informed search spaces.
    • Validity: Correctness across the entire question–method–execution–conclusion chain, including evidence-claim alignment, control adequacy, and statistical soundness.
    • Impact: Downstream scientific utility, measured through artifact adoption, code/benchmark reuse, and longitudinal follow-up.
    • Reliability: Consistency across reruns, robustness to seed and prompt perturbations, and system resilience under environment noise.
    • Provenance: Traceability of data sources, tool invocations, decision checkpoints, reviews, and human interventions.

    2. Autonomy Assessment Variables

    • Task Substitution: The specific workflow stages autonomously executed by the system versus those requiring human labor.
    • Decision Authority: The locus of control over branch selection, rejection, stopping, and revision.
    • Workflow Closure: The degree to which grounding, planning, execution, validation, and reporting form an unbroken loop.
    • Responsibility Retention: Allocation of liability and final accountability for errors, safety violations, and invalid claims.
  10. Knowl 10 — Domain-Conditioned Autonomy Ceilings across Scientific Disciplines

    theoretical result

    The achievable autonomy ceiling of AutoResearch is domain-conditioned, varying across disciplines based on artifact executability, physical manipulability, feedback latency, reversibility, and accountability burdens:

    • Computational and Formal Sciences (Highest Ceiling: Advanced L2L2): Core artifacts (code, datasets, proof scripts, benchmarks) are digital, replayable, and rapidly verifiable, supporting dense pipeline integration (L2-PL2\text{-P}) from ideation to drafting.
    • Physical Sciences, Chemistry, Materials, and Embodied AI (Intermediate Ceiling: L2-SL2\text{-S} to Bounded L2-PL2\text{-P}): Digital twins, numerical simulators, and robotic platforms close narrow experimental and calibration loops; however, broad autonomy is constrained by instrument drift, physical synthesis constraints, sim-to-real gaps, and apparatus variability.
    • Biology, Biomedicine, Medicine, Social Sciences, and Earth Sciences (Lowest Ceiling: Early-to-Middle L2L2): Autonomy is restricted to evidence synthesis, literature mining, and data curation. Full workflow closure is blocked by non-manipulable Earth systems, biological nonstationarity, patient safety risks, regulatory compliance, causal identification challenges, and institutional accountability requirements.
  11. Knowl 11 — Architectural Bottlenecks in Hypothesis Generation and Reflexive Iteration

    limitation

    Current LLM-based research pipelines exhibit two fundamental structural limitations that prevent genuine scientific agency:

    1. Combinatorial Hypothesis Formulation: Current agentic workflows decompose research into sequential template-matching steps. While effective for drafting, coding, and retrieval, hypothesis generation is predominantly restricted to superficial concept recombination of the form A+B→CA + B \to C (where AA and BB exist in prior literature), failing to perform true abductive reasoning from empirical anomalies to novel explanatory theories.
    2. Absence of Reflexive Iteration: Pipeline architectures operate as unidirectional pipelines where experimental results inform manuscript writing but do not back-propagate to revise underlying hypotheses, theoretical frameworks, or problem formulations. While tree-search methods explore implementation variants, they optimize within fixed research formulations rather than executing reflexive re-theorization upon finding negative or contradictory results.
  12. Knowl 12 — Reliability Compounding, Attack Surfaces, and Auditability in AutoResearch

    limitation

    Multi-stage AutoResearch systems face critical vulnerabilities across reliability, security, and auditability:

    • Hallucination Compounding: Early-stage LLM errors (e.g., hallucinated citations, misread papers, distorted baselines) are absorbed into intermediate artifacts (e.g., related work, experiment plans) and amplified across downstream coding, execution, and interpretation stages.
    • Workflow-Level Security Vulnerabilities: Modular research infrastructures introduce novel attack surfaces, including prompt injection embedded in retrieved scientific literature/tools and model-in-skill backdoor attacks (e.g., BadSkill) that corrupt reusable agent skills to execute hidden malicious behaviors during automated experiments.
    • Traceability vs. Auditability Gap: Mere operational logging of prompts and outputs is insufficient for auditability. Systems require structured provenance tracking, intermediate state checkpointing, and error localization mechanisms that explicitly separate model inference hallucinations, tool execution failures, orchestration bugs, and human intervention boundaries.

Coverage note — None was omitted; all key foundational concepts, the 5-level autonomy spectrum, the 5 workflow stages and their internal technical regimes, the dual-target evaluation framework, the domain-conditioned autonomy ceilings, and the major structural limitations and vulnerabilities have been synthesized into standalone knowls.

References

  1. 1.Ziming Luo, Zonglin Yang, Zexin Xu, Wei Yang, and Xinya Du. Llm4sr: A survey on large language models for scientific research. arXiv preprint arXiv:2501.04306, 2025.
  2. 2.John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. nature, 596(7873):583–589, 2021.
  3. 3.Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes, and Christina Mack. Agentic ai for scientific discovery: A survey of progress, challenges, and future directions. arXiv preprint arXiv:2503.08979, 2025.
  4. 4.Jiaqi Wei, Yuejin Yang, Xiang Zhang, Yuhan Chen, Xiang Zhuang, Zhangyang Gao, Dongzhan Zhou, Guangshuai Wang, Zhiqiang Gao, Juntai Cao, et al. From ai for science to agentic science: A survey on autonomous scientific discovery. arXiv preprint arXiv:2508.14111, 2025.
  5. 5.Haoxuan Zhang, Ruochi Li, Yang Zhang, Ting Xiao, Jiangping Chen, Junhua Ding, and Haihua Chen. The evolving role of large language models in scientific innovation: Evaluator, collaborator, and scientist. arXiv preprint arXiv:2507.11810, 2025.
  6. 6.Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Zihao Wang, and Yangqiu Song. From automation to autonomy: A survey on large language models in scientific discovery. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 17744–17761, 2025.
  7. 7.Valerie Fu. Ai for science: Opportunities, challenges, and future directions. Authorea Preprints, 2025.
  8. 8.Md Musaddaqul Hasib, Sumin Jo, Harsh Sinha, Jifeng Song, Arun Das, Zhentao Liu, Hugh Galloway, Huey Huang, Kexun Zhang, Shou-Jiang Gao, et al. A process-centric survey of ai for scientific discovery through the exhyte framework. 2025.
  9. 9.Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist: Towards fully automated open-ended scientific discovery. arXiv preprint arXiv:2408.06292, 2024.
  10. 10.Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search. arXiv preprint arXiv:2504.08066, 2025.
  11. 11.Patrick Tser Jern Kon, Jiachen Liu, Qiuyi Ding, Yiming Qiu, Zhenning Yang, Yibo Huang, Jayanth Srinivasa, Myungjin Lee, Mosharaf Chowdhury, and Ang Chen. Curie: Toward rigorous and automated scientific experimentation with ai agents. arXiv preprint arXiv:2502.16069, 2025.
  12. 12.Yingming Pu, Tao Lin, and Hongyu Chen. Piflow: Principle-aware scientific discovery with multi-agent collaboration. arXiv preprint arXiv:2505.15047, 2025.
  13. 13.Zhenzhen Zhuang, Jiandong Chen, Hongfeng Xu, Yuwen Jiang, and Jialiang Lin. Large language models for automated scholarly paper review: A survey. Information Fusion, 124:103332, 2025.
  14. 14.Chengwei Liu, Chong Wang, Jiayue Cao, Jingquan Ge, Kun Wang, Lyuye Zhang, Ming-Ming Cheng, Penghai Zhao, Tianlin Li, Xiaojun Jia, et al. A vision for auto research with llm agents. arXiv preprint arXiv:2504.18765, 2025.
  15. 15.Shubham Agarwal, Gaurav Sahu, Abhay Puri, Issam H Laradji, Krishnamurthy DJ Dvijotham, Jason Stanley, Laurent Charlin, and Christopher Pal. Litllm: A toolkit for scientific literature review. arXiv preprint arXiv:2402.01788, 2024.
  16. 16.GitHub repository. Openscholar. https://github.com/AkariAsai/OpenScholar, 2026.
  17. 17.Michael D Skarlinski, Sam Cox, Jon M Laurent, James D Braza, Michaela Hinks, Michael J Hammerling, Manvitha Ponnapati, Samuel G Rodriques, and Andrew D White. Language agents achieve superhuman synthesis of scientific knowledge. arXiv preprint arXiv:2409.13740, 2024.
  18. 18.GitHub repository. Paperqa2. https://github.com/Future-House/paper-qa, 2026.
  19. 19.Xingyao Wang, Boxuan Li, Yufan Song, Frank F Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al. Openhands: An open platform for ai software developers as generalist agents. In International Conference on Learning Representations, volume 2025, pages 65882–65919, 2025.
  20. 20.GitHub repository. Aider. https://github.com/Aider-AI/aider, 2026.
  21. 21.GitHub repository. Swe-agent. https://github.com/SWE-agent/SWE-agent, 2026.
  22. 22.GitHub repository. Agent laboratory. https://github.com/SamuelSchmidgall/AgentLaboratory, 2026.
  23. 23.GitHub repository. Ai-researcher. https://github.com/HKUDS/AI-Researcher, 2026.
  24. 24.GitHub repository. Auto-claude-code-research-in-sleep (aris). https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep, 2026.
  25. 25.OpenRaiser. Nanoresearch, 2026.
  26. 26.Alisia Lupidi, Bhavul Gauri, Thomas Simon Foster, Bassel Al Omari, Despoina Magka, Alberto Pepe, Alexis Audran-Reiss, Muna Aghamelu, Nicolas Baldwin, Lucia Cipolina-Kun, et al. Airs-bench: a suite of tasks for frontier ai research science agents. arXiv preprint arXiv:2602.06855, 2026.
  27. 27.Guijin Son, Jiwoo Hong, Honglu Fan, Heejeong Nam, Hyunwoo Ko, Seungwon Lim, Jinyeop Song, Jinha Choi, Gonçalo Paulo, Youngjae Yu, et al. When ai co-scientists fail: Spot-a benchmark for automated verification of scientific research. arXiv preprint arXiv:2505.11855, 2025.
  28. 28.Amal Gueroudji, Tanwi Mallick, Renan Souza, Rafael Ferreira Da Silva, Robert Ross, Matthieu Dorier, Philip Carns, Kyle Chard, and Ian Foster. Controla: Agentic workflow control mechanisms for reliable science. In 2025 IEEE International Conference on eScience (eScience), pages 415–426. IEEE, 2025.
  29. 29.Qiujie Xie, Yixuan Weng, Minjun Zhu, Fuchen Shen, Shulin Huang, Zhen Lin, Jiahui Zhou, Zilan Mao, Zijie Yang, Linyi Yang, et al. How far are ai scientists from changing the world? arXiv preprint arXiv:2507.23276, 2025.
  30. 30.Guiyao Tie, Pan Zhou, and Lichao Sun. A survey of ai scientists. arXiv preprint arXiv:2510.23045, 2025.
  31. 31.Qiguang Chen, Mingda Yang, Libo Qin, Jinhao Liu, Zheng Yan, Jiannan Guan, Dengyun Peng, Yiyan Ji, Hanjing Li, Mengkang Hu, et al. Ai4research: A survey of artificial intelligence for scientific research. arXiv preprint arXiv:2507.01903, 2025.
  32. 32.Karl Popper. The logic of scientific discovery. Routledge, 2005.
  33. 33.Thomas S Kuhn and Ian Hacking. The structure of scientific revolutions, volume 2. University of Chicago press Chicago, 1970.
  34. 34.Robert K Merton. The sociology of science: Theoretical and empirical investigations. University of Chicago press, 1973.
  35. 35.Jack Gallifant, Amelia Fiske, Yulia A Levites Strekalova, Juan S Osorio-Valencia, Rachael Parke, Rogers Mwavu, Nicole Martinez, Judy Wawira Gichoya, Marzyeh Ghassemi, Dina Demner-Fushman, et al. Peer review of gpt-4 technical report and systems card. PLOS digital health, 3(1):e0000417, 2024.
  36. 36.DeepSeek-AI. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.
  37. 37.GitHub repository. Openhands. https://github.com/All-Hands-AI/OpenHands, 2026.
  38. 38.Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, et al. Towards an ai co-scientist. arXiv preprint arXiv:2502.18864, 2025.
  39. 39.Ed Li, Junyu Ren, Xintian Pan, Cat Yan, Chuanhao Li, Dirk Bergemann, and Zhuoran Yang. Build your personalized research group: A multiagent framework for continual and interactive science automation. arXiv preprint arXiv:2510.15624, 2025.
  40. 40.Joeran Beel, Min-Yen Kan, and Moritz Baumgart. Evaluating sakana’s ai scientist: Bold claims, mixed results, and a promising future? In ACM SIGIR Forum, volume 59, pages 1–20. ACM New York, NY, USA, 2025.
  41. 41.Shreyansh Agrawal, Harsh B Anadkat, Kiran K Athimoolam, Harsh Bhardwaj, Trishul Chowdhury, Shengtao Gao, Purva K Kamat, Vishwadeepsinh Makwana, Mohammed H Shariff, Amitesh Badkul, et al. Can ai conduct autonomous scientific research? case studies on two real-world tasks. bioRxiv, pages 2026–01, 2026.
  42. 42.Ziming Luo, Atoosa Kasirzadeh, and Nihar B Shah. The more you automate, the less you see: Hidden pitfalls of ai scientist systems. arXiv preprint arXiv:2509.08713, 2025.
  43. 43.Alexander V Tobias and Adam Wahab. Autonomous ‘self-driving’laboratories: a review of technology and policy implications. Royal Society Open Science, 12(7):250646, 2025.
  44. 44.Shanghua Gao, Ada Fang, Yepeng Huang, Valentina Giunchiglia, Ayush Noori, Jonathan Richard Schwarz, Yasha Ektefaie, Jovana Kondic, and Marinka Zitnik. Empowering biomedical discovery with ai agents. Cell, 187(22):6125–6151, 2024.
  45. 45.Jiabin Tang, Lianghao Xia, Zhonghang Li, and Chao Huang. Ai-researcher: Autonomous scientific innovation. Advances in Neural Information Processing Systems, 38:9481–9520, 2026.
  46. 46.Kan Hatakeyama-Sato, Toshihiko Nishida, Kenta Kitamura, Yoshitaka Ushiku, Koichi Takahashi, Yuta Nabae, and Teruaki Hayakawa. Perspective on utilizing foundation models for laboratory automation in materials research. Science and Technology of Advanced Materials: Methods, 5(1):2582379, 2025.
  47. 47.Shuxiang Cao, Zijian Zhang, Mohammed Alghadeer, Simone D Fasciati, Michele Piscitelli, Mustafa Bakr, Peter Leek, and Alán Aspuru-Guzik. Agents for self-driving laboratories applied to quantum computing. arXiv preprint arXiv:2412.07978, 2024.
  48. 48.Rosni Vasu, Chandrayee Basu, Bhavana Dalvi, Cristina Sarasua, Peter Clark, and Abraham Bernstein. Hyper: Literature-grounded hypothesis generation and distillation with provenance. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 25424–25449, 2025.
  49. 49.Odhran O’Donoghue, Aleksandar Shtedritski, John Ginger, Ralph Abboud, Ali Ghareeb, and Samuel Rodriques. Bioplanner: automatic evaluation of llms on protocol planning in biology. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2676–2694, 2023.
  50. 50.Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, and Sung Ju Hwang. Researchagent: Iterative research idea generation over scientific literature with large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 6709–6738, 2025.
  51. 51.Chenyang Shao, Dehao Huang, Yu Li, Keyu Zhao, Weiquan Lin, Yining Zhang, Qingbin Zeng, Zhiyu Chen, Tianxing Li, Yifei Huang, et al. Omniscientist: Toward a co-evolving ecosystem of human and ai scientists. arXiv preprint arXiv:2511.16931, 2025.
  52. 52.Peter Jansen, Oyvind Tafjord, Marissa Radensky, Pao Siangliulue, Tom Hope, Bhavana Dalvi, Bodhisattwa Prasad Majumder, Daniel S Weld, and Peter Clark. Codescientist: End-to-end semi-automated scientific discovery with code-based experimentation. In Findings of the Association for Computational Linguistics: ACL 2025, pages 13370–13467, 2025.
  53. 53.Lianhao Zhou, Hongyi Ling, Cong Fu, Yepeng Huang, Michael Sun, Wendi Yu, Xiaoxuan Wang, Xiner Li, Xingyu Su, Junkai Zhang, et al. Autonomous agents for scientific discovery: Orchestrating scientists, language, code, and physics. arXiv preprint arXiv:2510.09901, 2025.
  54. 54.Tingting Chen, Srinivas Anumasa, Beibei Lin, Vedant Shah, Anirudh Goyal, and Dianbo Liu. Auto-bench: An automated benchmark for scientific discovery in llms. arXiv preprint arXiv:2502.15224, 2025.
  55. 55.Zifeng Wang, Benjamin Danek, and Jimeng Sun. Biodsa-1k: Benchmarking data science agents for biomedical research. arXiv preprint arXiv:2505.16100, 2025.
  56. 56.Yujie Liu, Zonglin Yang, Tong Xie, Jinjie Ni, Ben Gao, Yuqiang Li, Shixiang Tang, Wanli Ouyang, Erik Cambria, and Dongzhan Zhou. Researchbench: Benchmarking llms in scientific discovery via inspirationbased task decomposition. arXiv preprint arXiv:2503.21248, 2025.
  57. 57.Tianze Xu, Pengrui Lu, Lyumanshan Ye, Xiangkun Hu, and Pengfei Liu. Researcherbench: Evaluating deep ai research systems on the frontiers of scientific inquiry. arXiv preprint arXiv:2507.16280, 2025.
  58. 58.GitHub repository. autoresearch. https://github.com/karpathy/autoresearch, 2026.
  59. 59.GitHub repository. Deerflow. https://github.com/bytedance/deer-flow, 2026.
  60. 60.GitHub repository. Open deep research. https://github.com/langchain-ai/open_deep_research, 2026.
  61. 61.Stefan Kramer, Mattia Cerrato, Jannis Brugger, Sašo Džeroski, and Ross D King. Automated scientific discovery: from equation discovery to autonomous discovery systems. Machine Learning, 115(5):109, 2026.
  62. 62.Boyuan Zheng, Zerui Fang, Zhe Xu, Rui Wang, Yiwen Chen, Cunshi Wang, Mengwei Qu, Lei Lei, Zhen Feng, Yan Liu, et al. Agent4s: The transformation of research paradigms from the perspective of large language models. arXiv preprint arXiv:2506.23692, 2025.
  63. 63.Derek J De Solla Price. Little science, big science. Columbia university press, 1963.
  64. 64.Ross D King, Kenneth E Whelan, Ffion M Jones, Philip GK Reiser, Christopher H Bryant, Stephen H Muggleton, Douglas B Kell, and Stephen G Oliver. Functional genomic hypothesis generation and experimentation by a robot scientist. Nature, 427(6971):247–252, 2004.
  65. 65.Silviu-Marian Udrescu and Max Tegmark. Ai feynman: A physics-inspired method for symbolic regression. Science advances, 6(16):eaay2631, 2020.
  66. 66.GitHub repository. Storm. https://github.com/stanford-oval/storm, 2026.
  67. 67.Xiaofeng Shi, Qian Kou, Yuduo Li, Ning Tang, Jinxin Xie, Longbin Yu, Songjing Wang, and Hua Zhou. Scisage: A multi-agent framework for high-quality scientific survey generation. arXiv preprint arXiv:2506.12689, 2025.
  68. 68.Haiyuan Wan, Chen Yang, Junchi Yu, Meiqi Tu, Jiaxuan Lu, Di Yu, Jianbao Cao, Ben Gao, Jiaqing Xie, Aoran Wang, et al. Deep research arena: The first exam of llms’ research abilities via seminar-grounded tasks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 33341–33349, 2026.
  69. 69.GitHub repository. Gpt researcher. https://github.com/assafelovic/gpt-researcher, 2026.
  70. 70.GitHub repository. Tongyi deepresearch. https://github.com/Alibaba-NLP/DeepResearch, 2026.
  71. 71.Daniil A Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. Autonomous chemical research with large language models. Nature, 624(7992):570–578, 2023.
  72. 72.Nathan J Szymanski, Bernardus Rendy, Yuxing Fei, Rishi E Kumar, Tanjin He, David Milsted, Matthew J McDermott, Max Gallant, Ekin Dogus Cubuk, Amil Merchant, et al. An autonomous laboratory for the accelerated synthesis of inorganic materials. Nature, 624(7990):86, 2023.
  73. 73.Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang. Cycleresearcher: Improving automated research via automated review. In International Conference on Learning Representations, volume 2025, pages 3669–3709, 2025.
  74. 74.Ruochen Li, Teerth Patel, Qingyun Wang, and Xinya Du. Mlr-copilot: Autonomous machine learning research based on large language models agents. arXiv preprint arXiv:2408.14033, 2024.
  75. 75.Xu Yang, Xiao Yang, Shikai Fang, Yifei Zhang, Jian Wang, Bowen Xian, Qizheng Li, Jingyuan Li, Minrui Xu, Yuante Li, et al. R&d-agent: An llm-agent framework towards autonomous data science, 2025.
  76. 76.Zijun Liu, Kaiming Liu, Yiqi Zhu, Xuanyu Lei, Zonghan Yang, Zhenhe Zhang, Peng Li, and Yang Liu. Aigs: Generating science from ai-powered automated falsification. arXiv preprint arXiv:2411.11910, 2024.
  77. 77.Kyle Swanson, Wesley Wu, Nash L Bulaong, John E Pak, and James Zou. The virtual lab of ai agents designs new sars-cov-2 nanobodies. Nature, 646(8085):716–723, 2025.
  78. 78.Alireza Ghafarollahi and Markus J Buehler. Sciagents: automating scientific discovery through bioinspired multi-agent intelligent graph reasoning. Advanced Materials, 37(22):2413523, 2025.
  79. 79.Erzhuo Shao, Yifang Wang, Yifan Qian, Zhenyu Pan, Han Liu, and Dashun Wang. Sciscigpt: advancing human–ai collaboration in the science of science. Nature Computational Science, 6(3):301–315, 2026.
  80. 80.Ali Essam Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Caralyn J. Szostkiewicz, Dmytro Shved, Gavin J. Gyimesi, Jon M. Laurent, Samantha M. Wright, Muhammed T. Razzak, Andrew D. White, Silvia C. Finnemann, Michaela M. Hinks, and Samuel G. Rodriques. A multi-agent system for automating scientific discovery. Nature, May 2026.
  81. 81.Samuel Schmidgall and Michael Moor. Agentrxiv: Towards collaborative autonomous research. arXiv preprint arXiv:2503.18102, 2025.
  82. 82.Chen Zhu and Xiaolu Wang. Hler: Human-in-the-loop economic research via multi-agent pipelines for empirical discovery. arXiv preprint arXiv:2603.07444, 2026.
  83. 83.Dingjie Song, Hanrong Zhang, Dawei Liu, Yixin Liu, Zongxia Li, Zhengqing Yuan, Siqi Zhang, and Lichao Sun. Dr. claw: An ai research workspace from idea to paper, 2026.
  84. 84.Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Artiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, Anil Palepu, Keran Rong, Ryutaro Tanno, Khaled Saab, Fan Zhang, Jacob Blum, Andrew Carroll, Kavita Kulkarni, Nenad Tomašev, Dina Zverinski, Ivor Rendulic, Elahe Vedadi, Florian Hasler, Luka Rimanic, Marina Boia, Ivan Budiselic, Ben Feinstein, Mathias Bellaiche, Tom Sheffer, Jan Freyberg, Jeremy Ratcliff, Ottavia Bertolli, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Yuan Guan, Vikram Dhillon, Eeshit Dhaval Vaishnav, Byron Lee, Tiago R. D. Costa, José R. Penadés, Gary Peltz, Yossi Matias, James Manyika, Demis Hassabis, Yunhan Xu, Pushmeet Kohli, Annalisa Pawlosky, Alan Karthikesalingam, and Vivek Natarajan. Accelerating scientific discovery with co-scientist. Nature, May 2026.
  85. 85.GitHub repository. Idea2paper. https://github.com/AgentAlphaAGI/Idea2Paper, 2026.
  86. 86.Alexander Novikov, Ngân Vu, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco JR Ruiz, Abbas Mehrabian, et al. Alphaevolve: A coding agent for scientific and algorithmic discovery. arXiv preprint arXiv:2506.13131, 2025.
  87. 87.Yixuan Weng, Minjun Zhu, Qiujie Xie, Qiyao Sun, Zhen Lin, Sifan Liu, and Yue Zhang. Deepscientist: Advancing frontier-pushing scientific findings progressively. arXiv preprint arXiv:2509.26603, 2025.
  88. 88.Shiyang Feng, Runmin Ma, Xiangchao Yan, Yue Fan, Yusong Hu, Songtao Huang, Shuaiyu Zhang, Zongsheng Cao, Tianshuo Peng, Jiakang Yuan, et al. Internagent-1.5: A unified agentic framework for long-horizon autonomous scientific discovery. arXiv preprint arXiv:2602.08990, 2026.
  89. 89.Ludovico Mitchener, Angela Yiu, Benjamin Chang, Mathieu Bourdenx, Tyler Nadolski, Arvis Sulovari, Eric C Landsness, Daniel L Barabasi, Siddharth Narayanan, Nicky Evans, et al. Kosmos: An ai scientist for autonomous discovery. arXiv preprint arXiv:2511.02824, 2025.
  90. 90.GitHub repository. Researchclaw. https://github.com/ymx10086/ResearchClaw, 2026.
  91. 91.GitHub repository. Scienceclaw. https://github.com/beita6969/ScienceClaw, 2026.
  92. 92.Jiaqi Liu, Peng Xia, Siwei Han, Shi Qiu, Letian Zhang, Guiming Chen, Haoqin Tu, Xinyu Yang, Jiawei Zhou, Hongtu Zhu, Yun Li, Jiaheng Zhang, Yuyin Zhou, Zeyu Zheng, Cihang Xie, Mingyu Ding, and Huaxiu Yao. Autoresearchclaw: Fully autonomous research from idea to paper, 2026.
  93. 93.Yougang Lyu, Xi Zhang, Xinhao Yi, Yuyue Zhao, Shuyu Guo, Wenxiang Hu, Jan Piotrowski, Jakub Kaliski, Jacopo Urbani, Zaiqiao Meng, et al. Evoscientist: Towards multi-agent evolving ai scientists for end-to-end scientific discovery. arXiv preprint arXiv:2603.08127, 2026.
  94. 94.Cheng Wang, Zhibin He, Zhihao Peng, Shengyuan Liu, Yufan Hu, Lichao Sun, Xiang Li, and Yixuan Yuan. Neuroclaw technical report. arXiv preprint arXiv:2604.24696, 2026.
  95. 95.Eser Aygün, Anastasiya Belyaeva, Gheorghe Comanici, Marc Coram, Hao Cui, Jake Garrison, Renee Johnston, Anton Kast, Cory Y. McLean, Peter Norgaard, Zahra Shamsi, David Smalling, James Thompson, Subhashini Venugopalan, Brian P. Williams, Chujun He, Sarah Martinson, Martyna Plomecka, Lai Wei, Yuchen Zhou, Qian-Ze Zhu, Matthew Abraham, Erica Brand, Anna Bulanova, Jeffrey A. Cardille, Chris Co, Scott Ellsworth, Grace Joseph, Malcolm Kane, Ryan Krueger, Johan Kartiwa, Dan Liebling, Jan-Matthis Lueckmann, Paul Raccuglia, Xuefei Julie Wang, Katherine Chou, James Manyika, Yossi Matias, John C. Platt, Lizzie Dorfman, Shibl Mourad, and Michael P. Brenner. An ai system to help scientists write expert-level empirical software. Nature, May 2026.
  96. 96.Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, et al. Fire-bench: Evaluating agents on the rediscovery of scientific insights. arXiv preprint arXiv:2602.02905, 2026.
  97. 97.Tal Ifargan, Lukas Hafner, Maor Kern, Ori Alcalay, and Roy Kishony. Autonomous llm-driven research—from data to human-verifiable research papers. NEJM AI, 2(1):AIoa2400555, 2025.
  98. 98.Lukas Weidener, Marko Brkic, Mihailo Jovanovic, Ritvik Singh, Chiara Baccin, Emre Ulgac, Alex Dobrin, and Aakaash Meduri. Rethinking the ai scientist: Interactive multi-agent workflows for scientific discovery. arXiv preprint arXiv:2601.12542, 2026.
  99. 99.Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, and Emad Barsoum. Agent laboratory: Using llm agents as research assistants. Findings of the Association for Computational Linguistics: EMNLP 2025, pages 5977–6043, 2025.
  100. 100.Marissa Radensky, Simra Shahid, Raymond Fok, Pao Siangliulue, Tom Hope, and Daniel S Weld. Scideator: Human-llm scientific idea generation grounded in research-paper facet recombination. arXiv preprint arXiv:2409.14634, 2024.
  101. 101.NovelSeek Team, Bo Zhang, Shiyang Feng, Xiangchao Yan, Jiakang Yuan, Zhiyin Yu, Xiaohan He, Songtao Huang, Shaowei Hou, Zheng Nie, et al. Novelseek: When agent becomes the scientist–building closed-loop system from hypothesis to verification. arXiv e-prints, pages arXiv–2505, 2025.
  102. 102.Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. Chemcrow: Augmenting large-language models with chemistry tools. arXiv preprint arXiv:2304.05376, 2023.
  103. 103.Kexin Huang, Serena Zhang, Hanchen Wang, Yuanhao Qu, Yingzhou Lu, Yusuf Roohani, Ryan Li, Lin Qiu, Gavin Li, Junze Zhang, et al. Biomni: A general-purpose biomedical ai agent. biorxiv, 2025.
  104. 104.Fan Liu, Zherui Yang, Cancheng Liu, Tianrui Song, Xiaofeng Gao, and Hao Liu. Mm-agent: Llm as agents for real-world mathematical modeling problem. Advances in Neural Information Processing Systems, 38:20881–20934, 2026.
  105. 105.Dawei Li, Zongxia Li, Hongyang Du, Xiyang Wu, Shihang Gui, Yongbei Kuang, and Lichao Sun. Graph of skills: Dependency-aware structural retrieval for massive agent skills. arXiv preprint arXiv:2604.05333, 2026.
  106. 106.Benjamin Burger, Phillip M Maffettone, Vladimir V Gusev, Catherine M Aitchison, Yang Bai, Xiaoyan Wang, Xiaobo Li, Ben M Alston, Buyi Li, Rob Clowes, et al. A mobile robotic chemist. Nature, 583(7815):237–241, 2020.
  107. 107.Gihan Panapitiya, Emily Saldanha, Heather Job, and Olivia Hess. Autolabs: Cognitive multi-agent systems with self-correction for autonomous chemical experimentation. arXiv preprint arXiv:2509.25651, 2025.
  108. 108.Kourosh Darvish, Marta Skreta, Yuchi Zhao, Naruki Yoshikawa, Sagnik Som, Miroslav Bogdanovic, Yang Cao, Han Hao, Haoping Xu, Alán Aspuru-Guzik, et al. Organa: A robotic assistant for automated chemistry experimentation and characterization. Matter, 8(2), 2025.
  109. 109.Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, et al. Paperbench: Evaluating ai’s ability to replicate ai research. arXiv preprint arXiv:2504.01848, 2025.
  110. 110.Rui Li, Jia-Chen Gu, Po-Nien Kung, Heming Xia, Xiangwen Kong, Zhifang Sui, Nanyun Peng, et al. Llm-reval: Can we trust llm reviewers yet? arXiv preprint arXiv:2510.12367, 2025.
  111. 111.Weixin Liang, Yuhui Zhang, Hancheng Cao, Binglu Wang, Daisy Yi Ding, Xinyu Yang, Kailas Vodrahalli, Siyu He, Daniel Scott Smith, Yian Yin, et al. Can large language models provide useful feedback on research papers? a large-scale empirical analysis. NEJM AI, 1(8):AIoa2400196, 2024.
  112. 112.Birupaksha Biswas and Suhena Sarkar. Responsible agentic artificial intelligence governance: Risk, safety, and ethical challenges in autonomous systems. International Journal of Applied Resilience and Sustainability, 2(2):142–167, 2026.
  113. 113.Yujing Ke, Kevin George, Kathan Pandya, David Blumenthal, Maximilian Sprang, Gerrit Großmann, Sebastian Vollmer, and David Antony Selby. Biodisco: Multi-agent hypothesis generation with dual-mode evidence, iterative feedback and temporal evaluation. arXiv preprint arXiv:2508.01285, 2025.
  114. 114.Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers. In International Conference on Learning Representations, volume 2025, pages 94003–94092, 2025.
  115. 115.Patrick Tser Jern Kon, Jiachen Liu, Xinyi Zhu, Qiuyi Ding, Jingjia Peng, Jiarong Xing, Yibo Huang, Yiming Qiu, Jayanth Srinivasa, Myungjin Lee, et al. Exp-bench: Can ai conduct ai research experiments? arXiv preprint arXiv:2505.24785, 2025.
  116. 116.Yanzheng Xiang, Hanqi Yan, Shuyin Ouyang, Lin Gui, and Yulan He. Scireplicate-bench: Benchmarking llms in agent-driven algorithmic reproduction from research papers. arXiv preprint arXiv:2504.00255, 2025.
  117. 117.Ken Gu, Ruoxi Shang, Ruien Jiang, Keying Kuang, Richard-John Lin, Donghe Lyu, Yue Mao, Youran Pan, Teng Wu, Jiaqian Yu, et al. Blade: Benchmarking language model agents for data-driven science. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 13936–13971, 2024.
  118. 118.Jiayu Wang, Yifei Ming, Riya Dulepet, Qinglin Chen, Austin Xu, Zixuan Ke, Frederic Sala, Aws Albarghouthi, Caiming Xiong, and Shafiq Joty. Liveresearchbench: A live benchmark for user-centric deep research in the wild. arXiv preprint arXiv:2510.14240, 2025.
  119. 119.Qian Huang, Jian Vora, Percy Liang, and Jure Leskovec. Mlagentbench: Evaluating language agents on machine learning experimentation. arXiv preprint arXiv:2310.03302, 2023.
  120. 120.Ori Press, Andreas Hochlehnert, Ameya Prabhu, Vishaal Udandarao, Ofir Press, and Matthias Bethge. Citeme: Can language models accurately cite scientific claims? Advances in Neural Information Processing Systems, 37:7847–7877, 2024.
  121. 121.Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal, Danqi Chen, and Tianyu Gao. Litsearch: A retrieval benchmark for scientific literature search. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 15068–15083, 2024.
  122. 122.Leon Staufer, Kevin Feng, Kevin Wei, Luke Bailey, Yawen Duan, Mick Yang, A Pinar Ozisik, Stephen Casper, and Noam Kolt. The 2025 ai agent index: Documenting technical and safety features of deployed agentic ai systems. arXiv preprint arXiv:2602.17753, 2026.
  123. 123.Max Zimmer, Nico Pelleriti, Christophe Roux, and Sebastian Pokutta. The agentic researcher: A practical guide to ai-assisted research in mathematics and machine learning. arXiv preprint arXiv:2603.15914, 2026.
  124. 124.Bodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal, Bhavana Dalvi Mishra, Abhijeetsingh Meena, Aryan Prakhar, Tirth Vora, Tushar Khot, Ashish Sabharwal, and Peter Clark. Discoverybench: Towards datadriven discovery with large language models. In International Conference on Learning Representations, volume 2025, pages 4556–4579, 2025.
  125. 125.Zachary S Siegel, Sayash Kapoor, Nitya Nagdir, Benedikt Stroebl, and Arvind Narayanan. Core-bench: Fostering the credibility of published research through a computational reproducibility agent benchmark. arXiv preprint arXiv:2409.11363, 2024.
  126. 126.Ziru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li, Zeyi Liao, Chen Wei, Zitong Lu, et al. Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery. In International Conference on Learning Representations, volume 2025, pages 96934–96990, 2025.
  127. 127.Liana Patel, Negar Arabzadeh, Harshit Gupta, Ankita Sundar, Ion Stoica, Matei Zaharia, and Carlos Guestrin. Deepscholar-bench: A live benchmark and automated evaluation for generative research synthesis. arXiv preprint arXiv:2508.20033, 2025.
  128. 128.Nikos I Bosse, Jon Evans, Robert G Gambee, Daniel Hnyk, Peter Mühlbacher, Lawrence Phillips, Dan Schwarz, Jack Wildman, et al. Deep research bench: Evaluating ai web research agents. arXiv preprint arXiv:2506.06287, 2025.
  129. 129.Amirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol, Curtis Fox, Amrutha Varshini Ramesh, Étienne Marcotte, Xing Han Lù, Nicolas Chapados, Spandana Gella, Peter West, et al. Drbench: A realistic benchmark for enterprise deep research. arXiv preprint arXiv:2510.00172, 2025.
  130. 130.J Gregory Pauloski, Kyle Chard, and Ian Foster. Agentic discovery: Closing the loop with cooperative agents. Computer, 58(10):20–27, 2025.
  131. 131.Tingjia Miao, Jiawen Dai, Jingkun Liu, Jinxin Tan, Muhua Zhang, Wenkai Jin, Yuwen Du, Tian Jin, Xianghe Pang, Zexi Liu, et al. Physmaster: Building an autonomous ai physicist for theoretical and computational physics research. arXiv preprint arXiv:2512.19799, 2025.
  132. 132.Xueyang Zhou, Yihan Sun, Xijie Gong, Guiyao Tie, Pan Zhou, Lichao Sun, and Yongchao Chen. Embodiedclaw: Conversational workflow execution for embodied ai development. arXiv preprint arXiv:2604.13800, 2026.
  133. 133.Ruiying Li, Yunlang Zhou, YuYao Zhu, Kylin Chen, Jingyuan Wang, Sukai Wang, Kongtao Hu, Minhui Yu, Bowen Jiang, Zhan Su, et al. Roboclaw: An agentic framework for scalable long-horizon robotic tasks. arXiv preprint arXiv:2603.11558, 2026.
  134. 134.Michael Ahn, Debidatta Dwibedi, Chelsea Finn, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Karol Hausman, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, et al. Autort: Embodied foundation models for large scale orchestration of robotic agents. arXiv preprint arXiv:2401.12963, 2024.
  135. 135.Zhiyuan Zhou, Pranav Atreya, You Liang Tan, Karl Pertsch, and Sergey Levine. Autoeval: Autonomous evaluation of generalist robot manipulation policies in the real world. arXiv preprint arXiv:2503.24278, 2025.
  136. 136.Yao Mu, Tianxing Chen, Zanxin Chen, Shijia Peng, Zhiqian Lan, Zeyu Gao, Zhixuan Liang, Qiaojun Yu, Yude Zou, Mingkun Xu, et al. Robotwin: Dual-arm robot benchmark with generative digital twins. In Proceedings of the computer vision and pattern recognition conference, pages 27649–27660, 2025.
  137. 137.Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation. arXiv preprint arXiv:2506.18088, 2025.
  138. 138.Lirui Wang, Yiyang Ling, Zhecheng Yuan, Mohit Shridhar, Chen Bao, Yuzhe Qin, Bailin Wang, Huazhe Xu, and Xiaolong Wang. Gensim: Generating robotic simulation tasks via large language models. In International Conference on Learning Representations, volume 2024, pages 4890–4924, 2024.
  139. 139.Pu Hua, Minghuan Liu, Annabella Macaluso, Yunfeng Lin, Weinan Zhang, Huazhe Xu, and Lirui Wang. Gensim2: Scaling robot data generation with multi-modal and reasoning llms. In Conference on Robot Learning, pages 5030–5066, 2024.
  140. 140.Yufei Wang, Zhou Xian, Feng Chen, Tsun-Hsuan Wang, Yian Wang, Katerina Fragkiadaki, Zackory Erickson, David Held, and Chuang Gan. Robogen: Towards unleashing infinite data for automated robot learning via generative simulation. arXiv preprint arXiv:2311.01455, 2023.
  141. 141.Ajay Mandlekar, Soroush Nasiriany, Bowen Wen, Iretiayo Akinola, Yashraj Narang, Linxi Fan, Yuke Zhu, and Dieter Fox. Mimicgen: A data generation system for scalable robot learning using human demonstrations. arXiv preprint arXiv:2310.17596, 2023.
  142. 142.Caelan Reed Garrett, A. Mandlekar, Bowen Wen, and Dieter Fox. Skillmimicgen: Automated demonstration generation for efficient skill learning and deployment. In Conference on Robot Learning, pages 2750–2790, 2024.
  143. 143.Tianyuan Dai, Josiah Wong, Yunfan Jiang, Chen Wang, Cem Gokmen, Ruohan Zhang, Jiajun Wu, and Li Fei-Fei. Automated creation of digital cousins for robust policy learning. arXiv preprint arXiv:2410.07408, 2024.
  144. 144.Justin Yu, Letian Fu, Huang Huang, Karim El-Refai, Rares Andrei Ambrus, Richard Cheng, Muhammad Zubair Irshad, and Ken Goldberg. Real2render2real: Scaling robot data without dynamics simulation or robot hardware. arXiv preprint arXiv:2505.09601, 2025.
  145. 145.Xinhai Li, Jialin Li, Ziheng Zhang, Rui Zhang, Fan Jia, Tiancai Wang, Haoqiang Fan, Kuo-Kun Tseng, and Ruiping Wang. Robogsim: A real2sim2real robotic gaussian splatting simulator. arXiv preprint arXiv:2411.11839, 2024.
  146. 146.Xiaoshen Han, Minghuan Liu, Yilun Chen, Junqiu Yu, Xiaoyang Lyu, Yang Tian, Bolun Wang, Weinan Zhang, and Jiangmiao Pang. Re3 sim: Generating high-fidelity simulation data via 3d-photorealistic real-to-sim for robotic manipulation. arXiv preprint arXiv:2502.08645, 2025.
  147. 147.Yu Fang, Yue Yang, Xinghao Zhu, Kaiyuan Zheng, Gedas Bertasius, D. Szafir, and Mingyu Ding. Rebot: Scaling robot learning with real-to-sim-to-real robotic video synthesis. In IEEE/RSJ International Conference on Intelligent Robots and Systems, pages 11351–11358, 2025.
  148. 148.John Harwell and Maria Gini. Sierra: A modular framework for accelerating research and improving reproducibility. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 9111–9117. IEEE, 2023.
  149. 149.Qinjie Lin, Guo Ye, and Han Liu. Ems®: A massive computational experiment management system towards data-driven robotics. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 9068–9075. IEEE, 2023.
  150. 150.Sean Wilson, Paul Glotfelter, Siddharth Mayya, Gennaro Notomista, Yousef Emam, Xiaoyi Cai, and Magnus Egerstedt. The robotarium: Automation of a remotely accessible, multi-robot testbed. IEEE Robotics and Automation Letters, 6(2):2922–2929, 2021.
  151. 151.Stefan Bauer, Manuel Wüthrich, Felix Widmaier, Annika Buchholz, Sebastian Stark, Anirudh Goyal, Thomas Steinbrenner, Joel Akpo, Shruti Joshi, Vincent Berenz, et al. Real robot challenge: A robotics competition in the cloud. In NeurIPS 2021 Competitions and Demonstrations Track, pages 190–204. PMLR, 2022.
  152. 152.Qing Zhu, Fei Zhang, Yan Huang, Hengyu Xiao, LuYuan Zhao, XuChun Zhang, Tao Song, XinSheng Tang, Xiang Li, Guo He, et al. An all-round ai-chemist with a scientific mind. National Science Review, 9(10):nwac190, 2022.
  153. 153.Amil Merchant, Simon Batzner, Samuel S Schoenholz, Muratahan Aykol, Gowoon Cheon, and Ekin Dogus Cubuk. Scaling deep learning for materials discovery. Nature, 624(7990):80–85, 2023.
  154. 154.Yixiang Ruan, Chenyin Lu, Ning Xu, Yuchen He, Yixin Chen, Jian Zhang, Jun Xuan, Jianzhang Pan, Qun Fang, Hanyu Gao, et al. An automatic end-to-end chemical synthesis development platform powered by large language models. Nature communications, 15(1):10160, 2024.
  155. 155.Tao Song, Man Luo, Xiaolong Zhang, Linjiang Chen, Yan Huang, Jiaqi Cao, Qing Zhu, Daobin Liu, Baicheng Zhang, Gang Zou, et al. A multiagent-driven robotic ai chemist enabling autonomous chemical research on demand. Journal of the American Chemical Society, 147(15):12534–12545, 2025.
  156. 156.Loïc M Roch, Florian Häse, and Alán Aspuru-Guzik. Chemos: An orchestration software to democratize autonomous discovery. 2020.
  157. 157.Samuel Alber, Bowen Chen, Eric Sun, Alina Isakova, Aaron J Wilk, and James Zou. Cellvoyager: Ai compbio agent generates new insights by autonomously analyzing biological data. Nature Methods, pages 1–11, 2026.
  158. 158.Mohammad HamediRad, Ran Chao, Scott Weisberg, Jiazhang Lian, Saurabh Sinha, and Huimin Zhao. Towards a fully automated algorithm driven platform for biosystems design. Nature communications, 10(1):5150, 2019.
  159. 159.Ievgeniia A Tiukova, Daniel Brunnsåker, Erik Y Bjurström, Alexander H Gower, Filip Kronström, Gabriel K Reder, Ronald S Reiserer, Konstantin Korovin, Larisa B Soldatova, John P Wikswo, et al. Genesis: towards the automation of systems biology research. arXiv preprint arXiv:2408.10689, 2024.
  160. 160.Yibo Qiu, Zan Huang, Zhiyu Wang, Handi Liu, Yiling Qiao, Yifeng Hu, Shu’ang Sun, Hangke Peng, Ronald X Xu, and Mingzhai Sun. Biomars: A multi-agent robotic system for autonomous biological experiments. arXiv preprint arXiv:2507.01485, 2025.
  161. 161.Mingyu Wu, Zhaoguo Wang, Jiabin Wang, Zhiyuan Dong, Jingkai Yang, Qingting Li, Tianyu Huang, Lei Zhao, Mingqiang Li, Fei Wang, et al. An ai-native experimental laboratory for autonomous biomolecular engineering. arXiv preprint arXiv:2507.02379, 2025.
  162. 162.Ariel Yuhan Ong, David A Merle, Siegfried K Wagner, and Pearse A Keane. Exploring the dilemma of ai use in medical research and knowledge synthesis: A perspective on deep research tools. Journal of medical Internet research, 27:e75666, 2025.
  163. 163.Achilleas Livieratos, Maria Kudela, Yuxi Zhao, All-shine Chen, Xin Luo, Junjing Lin, Di Zhang, Sai Dharmarajan, Sotirios Tsiodras, Vivek Rudrapatna, et al. Metamind: A multi-agent transformer-driven framework for automated network meta-analyses. Plos one, 21(2):e0342895, 2026.
  164. 164.Kaitlyn Hair, Emma Wilson, Charis Wong, Anthony Tsang, Malcolm Macleod, and Alexandra Bannach-Brown. Systematic online living evidence summaries: emerging tools to accelerate evidence synthesis. Clinical Science, 137(10):773–784, 2023.
  165. 165.Kyeryoung Lee, Surabhi Datta, Hunki Paek, Majid Rastegar-Mojarad, Liang-Chin Huang, Long He, Siwei Wang, Jingqi Wang, and Xiaoyan Wang. Aid-slr: A generative artificial intelligence-driven automated system for systematic literature review. MedRxiv, pages 2024–07, 2024.
  166. 166.Hongtao Wu, Boyun Zheng, Dingjie Song, Yu Jiang, Jianfeng Gao, Lei Xing, Lichao Sun, and Yixuan Yuan. Towards a medical ai scientist. arXiv preprint arXiv:2603.28589, 2026.
  167. 167.Jianing Qiu, Kyle Lam, Guohao Li, Amish Acharya, Tien Yin Wong, Ara Darzi, Wu Yuan, and Eric J Topol. Llm-based agentic systems in medicine and healthcare. Nature Machine Intelligence, 6(12):1418–1420, 2024.
  168. 168.Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, Jeffrey Jopling, and Michael Moor. Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments. arXiv preprint arXiv:2405.07960, 2024.
  169. 169.Keyu Zhao, Fengli Xu, Yong Li, and Tie-Yan Liu. Hybridquestion: Human-ai collaboration for identifying high-impact research questions. arXiv preprint arXiv:2602.03849, 2025.
  170. 170.Zijie Guo, Jiong Wang, Fenghua Ling, Wangxu Wei, Xiaoyu Yue, Zhe Jiang, Wanghan Xu, Jing-Jia Luo, Lijing Cheng, Yoo-Geun Ham, et al. A self-evolving ai agent system for climate science. arXiv preprint arXiv:2507.17311, 2025.
  171. 171.Kaikai Zhang, Xiang Wang, Haoluo Zhao, Nan Chen, Mengyang Yu Jing-Jia Luo, Tao Song, and Fan Meng. Tianji: An autonomous ai meteorologist for discovering physical mechanisms in atmospheric science. arXiv preprint arXiv:2603.27738, 2026.
  172. 172.Ahmed Jaber, Wangshu Zhu, Ayon Roy, Karthick Jayavelu, Justin Downes, Sameer Mohamed, Candace Agonafir, Linnia Hawkins, and Tian Zheng. Autoclimds: Climate data science agentic ai–a knowledge graph is all you need. arXiv preprint arXiv:2509.21553, 2025.
  173. 173.Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberger, Meire Fortunato, Ferran Alet, Suman Ravuri, Timo Ewalds, Zach Eaton-Rosen, Weihua Hu, et al. Learning skillful medium-range global weather forecasting. Science, 382(6677):1416–1421, 2023.
  174. 174.Ilan Price, Alvaro Sanchez-Gonzalez, Ferran Alet, Tom R Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, et al. Probabilistic weather forecasting with machine learning. Nature, 637(8044):84–90, 2025.
  175. 175.Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. Accurate medium-range global weather forecasting with 3d neural networks. Nature, 619(7970):533–538, 2023.
  176. 176.Tung Nguyen, Johannes Brandstetter, Ashish Kapoor, Jayesh K Gupta, and Aditya Grover. Climax: A foundation model for weather and climate. arXiv preprint arXiv:2301.10343, 2023.
  177. 177.Kevin J Boudreau, Eva C Guinan, Karim R Lakhani, and Christoph Riedl. Looking across and looking beyond the knowledge frontier: Intellectual distance, novelty, and resource allocation in science. Management science, 62(10):2765–2783, 2016.
  178. 178.Dashun Wang, Chaoming Song, and Albert-László Barabási. Quantifying long-term scientific impact. Science, 342(6154):127–132, 2013.
  179. 179.Michael Park, Erin Leahey, and Russell J Funk. Papers and patents are becoming less disruptive over time. Nature, 613(7942):138–144, 2023.
  180. 180.Guiyao Tie, Jiawen Shi, Pan Zhou, and Lichao Sun. Badskill: Backdoor attacks on agent skills via model-in-skill poisoning. arXiv preprint arXiv:2604.09378, 2026.

Citation

MLA
Tie, G., et al. “AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery”. arXiv, 2026, https://doi.org/10.48550/arxiv.2605.23204.
APA
Tie, G., Shi, J., Song, D., Huang, Y., Sheng, Z., Zhou, X., Liu, D., Zhou, P., Chen, Y., Xu, R., He, L., Wen, Q., Li, M., Lu, C., Li, S., Xie, P., Yuan, Y., Meng, R., Xing, L., … Gao, J. (2026). AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery. arXiv. https://doi.org/10.48550/arxiv.2605.23204
Chicago
Tie, G., J. Shi, D. Song, et al. 2026. “AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2605.23204.
Harvard
Tie, G. et al. (2026) “AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery”. arXiv. Available at: https://doi.org/10.48550/arxiv.2605.23204.
Vancouver
1. Tie G, Shi J, Song D, et al (2026) AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery. https://doi.org/10.48550/arxiv.2605.23204

BibTeX

@misc{https://doi.org/10.48550/arxiv.2605.23204,
  doi = {10.48550/ARXIV.2605.23204},
  url = {https://arxiv.org/abs/2605.23204},
  author = {Tie, Guiyao and Shi, Jiawen and Song, Dingjie and Huang, Yixiao and Sheng, Ziji and Zhou, Xueyang and Liu, Daizong and Zhou, Pan and Chen, Yongchao and Xu, Ran and He, Lifang and Wen, Qingsong and Li, Manling and Lu, Cong and Li, Shuai and Xie, Pengtao and Yuan, Yixuan and Meng, Rui and Xing, Lei and Sun, Lichao and Xiong, Caiming and Yu, Philip S. and Gao, Jianfeng},
  keywords = {Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/