Agentic Large Language Models, a Survey
Aske PlaatMax J. van DuijnNiki van SteinMike PreussPeter van der PuttenKees Joost Batenburg
Categorizes autonomous language models across reasoning, acting, and social interaction to explain how agentic behaviors generate new training data and overcome data scaling limits.
Recent advances in artificial intelligence face significant bottlenecks: conventional large language models struggle with multi-step reasoning, frequently produce factually incorrect outputs, and risk plateauing as usable static training data becomes scarce. At the same time, expanding models from passive text completion tools into autonomous decision-makers is essential for complex real-world workflows. To address these limitations, the article reviews the rapid development of agentic systems that actively engage with environments and generate their own empirical training data.
The main objective of the article is to provide a comprehensive survey and structured taxonomy of agentic large language models and outline a research agenda for their development. It evaluates how integrating reasoning, action execution, and social interaction transforms passive models into autonomous, goal-directed agents.
To conduct this evaluation, the article synthesizes recent high-impact literature across natural language processing, reinforcement learning, robotics, and multi-agent systems, primarily focusing on peer-reviewed and reputable studies from 2023 through 2025. It categorizes the body of work into three mutually reinforcing dimensions: reasoning mechanisms, acting capabilities, and multi-agent interactions.
The findings highlight five core insights across these dimensions. First, structured multi-step reasoning techniques, such as step-by-step prompting and external search trees, dramatically improve complex problem-solving; for instance, combining symbolic interpreters with code-based prompting raised mathematical problem-solving accuracy from 78.7% to 92.5%. Second, self-reflection loops and reinforcement learning allow models to evaluate their own intermediate outputs, reducing reasoning errors and generating synthetic training traces directly at inference time. Third, equipping models with external tools, application interfaces, and robotic control modules enables effective real-world automation in domains such as software engineering, financial analysis, and clinical workflows, where models sometimes exceed human performance in diagnostic accuracy. Fourth, collaborative multi-agent frameworks consistently outperform single monolithic prompts, with collaborative coding architectures improving benchmark task success rates from roughly 30% to 50%. Fifth, large-scale multi-agent simulations can spontaneously generate complex social phenomena, including shared norms, conventions, and collective coordination, without explicit role engineering.
These findings indicate that agency creates a self-sustaining cycle where reasoning, acting, and interacting continuously generate new grounded data, mitigating the risk of data depletion and reducing model hallucination. Operationally, these systems can lower costs and accelerate productivity in high-value sectors such as logistics, healthcare, and software development. However, autonomous action in physical and financial environments introduces operational, legal, and security risks, particularly when models encounter adversarial prompts or face unclear liability for critical errors.
Decision-makers should selectively pilot agentic workflows in lower-risk domains—such as research synthesis, code generation, and internal scheduling—while maintaining strict human-in-the-loop oversight for high-stakes medical and financial tasks. Organizations must also implement standard communication protocols and rigorous safety guardrails against prompt injection and tool misuse before deploying fully autonomous systems.
These conclusions should be interpreted with caution due to existing technical limitations. Current agentic models continue to struggle with abstract causal reasoning, spatial navigation, and long-horizon planning in partially observable environments, and recursive multi-agent loops remain susceptible to instability and conversational degradation. As the field matures, confidence is highest in structured, single-domain tool use and collaborative reasoning, whereas fully open-ended agent autonomy requires further empirical validation and robust governance frameworks.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). This 2023 survey establishes an early taxonomy and architecture for LLM-based agents, giving useful context for how the source organizes the field around reasoning, action, and interaction.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Its 2023 synthesis of agent architectures, capabilities, applications, and evaluation supplies foundational concepts that the source revisits in its broader survey.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). ReAct’s interleaving of reasoning and environment actions is a foundational agent method that clarifies the source’s reasoning-and-action categories.
- Paper: AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, Qingyun Wu et al. (2023). AutoGen’s early framework for role-based multi-agent conversation provides a concrete foundation for the collaborative systems surveyed in the source.
- Paper: Generative Agents: Interactive Simulacra of Human Behavior, Joon Sung Park et al. (2023). Generative Agents introduces memory, reflection, and planning in an interactive social simulation, grounding the source’s account of agent interaction and emergent behavior.
- Paper: CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Guohao Li et al. (2023). CAMEL’s role-playing agents offer an early example of autonomous multi-agent collaboration and interaction that informs the source’s treatment of agent societies.
- Paper: ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models, Jinheon Baek et al. (2025). ResearchAgent carries the survey’s account of agentic research into an evaluated system that generates and iteratively refines scientific proposals.
- Paper: Towards a Science of Scaling Agent Systems, Yubin Kim et al. (2025). This controlled study extends the survey’s discussion of multi-agent collaboration by measuring when agent teams help, and when coordination costs undermine them.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). Agent Workflow Memory develops the survey’s account of memory and learning into reusable procedures that improve later agent task execution.
- Paper: Magentic-UI: Towards Human-in-the-loop Agentic Systems, Hussein Mozannar et al. (2025). Magentic-UI advances the survey’s safety concerns into a tested human-in-the-loop design for overseeing agents that take actions across digital environments.
