On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
Jen-tse HuangJiaxu ZhouTailin JinXuhui ZhouZixi ChenWenxuan WangYouliang YuanMichael R. LyuMaarten Sap
Reveals how faulty agents degrade multi-agent LLM systems across different collaboration topologies, establishing that hierarchical workflows maintain the highest stability while introducing automated error injection methods and verification strategies that recover up to 96.4% of compromised performance.
Multi-agent systems powered by large language models have become widely adopted for complex tasks, yet they are increasingly vulnerable to operational disruptions when individual agents behave incompetently or introduce stealthy, malicious errors. This real-world vulnerability poses significant risks as organizations deploy decentralized, multi-agent frameworks across critical software, analytical, and translation workflows.
The article evaluates how different organizational structures and task types withstand faulty agents, and demonstrates practical defense mechanisms to restore system performance. The authors evaluate six systems representing three organizational models—linear pipelines, flat peer groups, and hierarchical setups—across four standard benchmarks covering code generation, mathematical problem solving, translation, and text evaluation. To rigorously test these setups, the researchers introduced two automated fault-generation methods: rewriting agent profiles to create stealthily faulty roles, and directly injecting controlled errors into intermediate communication messages.
The analysis reveals that organizational structure is the most decisive factor in determining resilience. Hierarchical structures demonstrated superior fault tolerance, suffering an average performance drop of only 5.5%, compared to drops of 10.5% in flat systems and 23.7% in linear pipelines. Tasks requiring strict formal logic and objectivity proved far more vulnerable; code generation suffered the steepest average performance decline at 22.6%, whereas subjective tasks such as translation experienced minor drops of under 5%. Additionally, subtle semantic errors and higher message failure frequencies degraded performance far more severely than syntax errors, as agents frequently overlook plausibly framed but logically flawed inputs. Introducing errors in high-level roles, such as planners or project managers, triggered cascading failures that disrupted the entire team.
These findings indicate that architectural choices directly govern operational risk and accuracy in multi-agent deployments. Hierarchical oversight allows central decision-makers to evaluate competing proposals, mitigating single points of failure. In contrast, linear workflows lack oversight, allowing errors to compound undetected. Interestingly, deliberate perturbations occasionally aided debate-based systems by breaking repetitive reasoning loops, highlighting that cognitive diversity can benefit systems when properly managed.
Organizations developing or deploying multi-agent systems should adopt hierarchical architectures with centralized coordination rather than sequential or purely flat arrangements. To actively counter agent faults, teams should implement defensive mechanisms: adding validation instructions that empower agents to challenge peer outputs, and deploying dedicated review agents to inspect and correct intermediate messages. Combining these defensive strategies restored up to 96.4% of lost performance in vulnerable systems.
Confidence in these structural insights is supported across multiple underlying language models, though readers should note the boundary conditions. The experiments were conducted on text-only benchmarks using automated evaluation metrics with a single faulty agent introduced at a time. Teams deploying multi-agent systems in multimodal environments, long-horizon planning scenarios, or complex real-world workflows should run targeted pilot tests to confirm resilience before full operational rollout.
- Paper: Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication, Zhangyue Yin et al. (2023). Its comparison of communication structures—including relay and hierarchical debate—provides a foundation for understanding how collaboration topology can shape the source paper’s resilience results.
- Paper: Improving Factuality and Reasoning in Language Models through Multiagent Debate, Yilun Du et al. (2024). It establishes multi-agent debate as a method for improving reasoning through critique, the collaborative mechanism whose reliability under faulty participation the source examines.
- Paper: Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate, Tian Liang et al. (2024). Its debate framework and evidence on independent judging help frame the source’s investigation of when multi-agent discussion can detect and correct flawed contributions.
- Paper: AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security, Dongrui Liu et al. (2026). It extends the source’s focus on detecting and containing agent failures into a diagnostic guardrail framework that identifies risky actions and traces their causes across agent trajectories.
