Agentic Large Language Models, a Survey

Aske PlaatMax J. van DuijnNiki van SteinMike PreussPeter van der PuttenKees Joost Batenburg

article2025JAIR198 citations

Categorizes autonomous language models across reasoning, acting, and social interaction to explain how agentic behaviors generate new training data and overcome data scaling limits.

Listen

Recent advances in artificial intelligence face significant bottlenecks: conventional large language models struggle with multi-step reasoning, frequently produce factually incorrect outputs, and risk plateauing as usable static training data becomes scarce. At the same time, expanding models from passive text completion tools into autonomous decision-makers is essential for complex real-world workflows. To address these limitations, the article reviews the rapid development of agentic systems that actively engage with environments and generate their own empirical training data.

The main objective of the article is to provide a comprehensive survey and structured taxonomy of agentic large language models and outline a research agenda for their development. It evaluates how integrating reasoning, action execution, and social interaction transforms passive models into autonomous, goal-directed agents.

To conduct this evaluation, the article synthesizes recent high-impact literature across natural language processing, reinforcement learning, robotics, and multi-agent systems, primarily focusing on peer-reviewed and reputable studies from 2023 through 2025. It categorizes the body of work into three mutually reinforcing dimensions: reasoning mechanisms, acting capabilities, and multi-agent interactions.

The findings highlight five core insights across these dimensions. First, structured multi-step reasoning techniques, such as step-by-step prompting and external search trees, dramatically improve complex problem-solving; for instance, combining symbolic interpreters with code-based prompting raised mathematical problem-solving accuracy from 78.7% to 92.5%. Second, self-reflection loops and reinforcement learning allow models to evaluate their own intermediate outputs, reducing reasoning errors and generating synthetic training traces directly at inference time. Third, equipping models with external tools, application interfaces, and robotic control modules enables effective real-world automation in domains such as software engineering, financial analysis, and clinical workflows, where models sometimes exceed human performance in diagnostic accuracy. Fourth, collaborative multi-agent frameworks consistently outperform single monolithic prompts, with collaborative coding architectures improving benchmark task success rates from roughly 30% to 50%. Fifth, large-scale multi-agent simulations can spontaneously generate complex social phenomena, including shared norms, conventions, and collective coordination, without explicit role engineering.

These findings indicate that agency creates a self-sustaining cycle where reasoning, acting, and interacting continuously generate new grounded data, mitigating the risk of data depletion and reducing model hallucination. Operationally, these systems can lower costs and accelerate productivity in high-value sectors such as logistics, healthcare, and software development. However, autonomous action in physical and financial environments introduces operational, legal, and security risks, particularly when models encounter adversarial prompts or face unclear liability for critical errors.

Decision-makers should selectively pilot agentic workflows in lower-risk domains—such as research synthesis, code generation, and internal scheduling—while maintaining strict human-in-the-loop oversight for high-stakes medical and financial tasks. Organizations must also implement standard communication protocols and rigorous safety guardrails against prompt injection and tool misuse before deploying fully autonomous systems.

These conclusions should be interpreted with caution due to existing technical limitations. Current agentic models continue to struggle with abstract causal reasoning, spatial navigation, and long-horizon planning in partially observable environments, and recursive multi-agent loops remain susceptible to instability and conversational degradation. As the field matures, confidence is highest in structured, single-domain tool use and collaborative reasoning, whereas fully open-ended agent autonomy requires further empirical validation and robust governance frameworks.

Cover for Agentic Large Language Models, a Survey

Abstract

Background: There is great interest in agentic LLMs, large language models that act as agents. Objectives: We review the growing body of work in this area and provide a research agenda. Methods: Agentic LLMs are LLMs that (1) reason, (2) act, and (3) interact. We organize the literature according to these three categories. Results: The research in the first category focuses on reasoning, reflection, and retrieval, aiming to improve decision making; the second category focuses on action models, robots, and tools, aiming for agents that act as useful assistants; the third category focuses on multi-agent systems, aiming for collaborative task solving and simulating interaction to study emergent social behavior. We find that works mutually benefit from results in other categories: retrieval enables tool use, reflection improves multi-agent collaboration, and reasoning benefits all categories. Conclusions: We discuss applications of agentic LLMs and provide an agenda for further research. Important applications are in medical diagnosis, logistics and financial market analysis. Meanwhile, self-reflective agents playing roles and interacting with one another augment the process of scientific research itself. Further, agentic LLMs provide a solution for the problem of LLMs running out of training data: inference-time behavior generates new training states, such that LLMs can keep learning without needing ever larger datasets. We note that there is risk associated with LLM assistants taking action in the real world—safety, liability and security are open problems—while agentic LLMs are also likely to benefit society.

Citation

MLA
Plaat, A., et al. “Agentic Large Language Models, a Survey”. Journal of Artificial Intelligence Research, vol. 84, 2025, https://doi.org/10.1613/jair.1.18675.
APA
Plaat, A., Van Duijn, M., Van Stein, N., Preuss, M., Van der Putten, P., & Batenburg, K. J. (2025). Agentic Large Language Models, a Survey. Journal of Artificial Intelligence Research, 84. https://doi.org/10.1613/jair.1.18675
Chicago
Plaat, A., M. Van Duijn, N. Van Stein, M. Preuss, P. Van der Putten, and K. J. Batenburg. 2025. “Agentic Large Language Models, a Survey”. Journal of Artificial Intelligence Research 84. https://doi.org/10.1613/jair.1.18675.
Harvard
Plaat, A. et al. (2025) “Agentic Large Language Models, a Survey”, Journal of Artificial Intelligence Research, 84. Available at: https://doi.org/10.1613/jair.1.18675.
Vancouver
1. Plaat A, Van Duijn M, Van Stein N, Preuss M, Van der Putten P, Batenburg KJ (2025) Agentic Large Language Models, a Survey. Journal of Artificial Intelligence Research. https://doi.org/10.1613/jair.1.18675

BibTeX

@article{Plaat_2025, title={Agentic Large Language Models, a Survey}, volume={84}, ISSN={1076-9757}, url={http://dx.doi.org/10.1613/jair.1.18675}, DOI={10.1613/jair.1.18675}, journal={Journal of Artificial Intelligence Research}, publisher={AI Access Foundation}, author={Plaat, Aske and Van Duijn, Max and Van Stein, Niki and Preuss, Mike and Van der Putten, Peter and Batenburg, Kees Joost}, year={2025}, month=Dec }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/