On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents

Jen-tse HuangJiaxu ZhouTailin JinXuhui ZhouZixi ChenWenxuan WangYouliang YuanMichael R. LyuMaarten Sap

article2025ICML122 citations

Reveals how faulty agents degrade multi-agent LLM systems across different collaboration topologies, establishing that hierarchical workflows maintain the highest stability while introducing automated error injection methods and verification strategies that recover up to 96.4% of compromised performance.

Listen

Multi-agent systems powered by large language models have become widely adopted for complex tasks, yet they are increasingly vulnerable to operational disruptions when individual agents behave incompetently or introduce stealthy, malicious errors. This real-world vulnerability poses significant risks as organizations deploy decentralized, multi-agent frameworks across critical software, analytical, and translation workflows.

The article evaluates how different organizational structures and task types withstand faulty agents, and demonstrates practical defense mechanisms to restore system performance. The authors evaluate six systems representing three organizational models—linear pipelines, flat peer groups, and hierarchical setups—across four standard benchmarks covering code generation, mathematical problem solving, translation, and text evaluation. To rigorously test these setups, the researchers introduced two automated fault-generation methods: rewriting agent profiles to create stealthily faulty roles, and directly injecting controlled errors into intermediate communication messages.

The analysis reveals that organizational structure is the most decisive factor in determining resilience. Hierarchical structures demonstrated superior fault tolerance, suffering an average performance drop of only 5.5%, compared to drops of 10.5% in flat systems and 23.7% in linear pipelines. Tasks requiring strict formal logic and objectivity proved far more vulnerable; code generation suffered the steepest average performance decline at 22.6%, whereas subjective tasks such as translation experienced minor drops of under 5%. Additionally, subtle semantic errors and higher message failure frequencies degraded performance far more severely than syntax errors, as agents frequently overlook plausibly framed but logically flawed inputs. Introducing errors in high-level roles, such as planners or project managers, triggered cascading failures that disrupted the entire team.

These findings indicate that architectural choices directly govern operational risk and accuracy in multi-agent deployments. Hierarchical oversight allows central decision-makers to evaluate competing proposals, mitigating single points of failure. In contrast, linear workflows lack oversight, allowing errors to compound undetected. Interestingly, deliberate perturbations occasionally aided debate-based systems by breaking repetitive reasoning loops, highlighting that cognitive diversity can benefit systems when properly managed.

Organizations developing or deploying multi-agent systems should adopt hierarchical architectures with centralized coordination rather than sequential or purely flat arrangements. To actively counter agent faults, teams should implement defensive mechanisms: adding validation instructions that empower agents to challenge peer outputs, and deploying dedicated review agents to inspect and correct intermediate messages. Combining these defensive strategies restored up to 96.4% of lost performance in vulnerable systems.

Confidence in these structural insights is supported across multiple underlying language models, though readers should note the boundary conditions. The experiments were conducted on text-only benchmarks using automated evaluation metrics with a single faulty agent introduced at a time. Teams deploying multi-agent systems in multimodal environments, long-horizon planning scenarios, or complex real-world workflows should run targeted pilot tests to confirm resilience before full operational rollout.

Cover for On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents

Abstract

Large language model-based multi-agent systems have shown great abilities across various tasks due to the collaboration of expert agents, each focusing on a specific domain. However, the impact of clumsy or even malicious agents—those who frequently make errors in their tasks—on the overall performance of the system remains under-explored. This paper investigates: (1) What is the resilience of various system structures (e.g., A→B→C, A↔B↔C) under faulty agents, on different downstream tasks? (2) How can we increase system resilience to defend against these agents? To simulate faulty agents, we propose two approaches—AutoTransform and AutoInject—which introduce mistakes into the agents’ responses. Experiments on four downstream tasks using six systems show that the “hierarchical” structure, i.e., A→(B↔C), exhibits superior resilience with the lowest performance drop of 5.5%, compared to 10.5% and 23.7% of other two structures. To further improve resilience, we introduce (1) Challenger, that introduces a mechanism for each agent to challenge others’ outputs, and (2) Inspector, an additional agent to review and correct messages, recovering up to 96.4% errors made by faulty agents. Our code and data are available at https://github.com/CUHK-ARISE/MAS-Resilience.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 3. Methodology: Introducing Errors
  • 3.1. AUTOTRANSFORM: Faulty Agent Transformation
  • 3.2. AUTOINJECT: Direct Error Injection
  • 4. Experiments
  • 4.1. RQ1: Impact of System Structures
  • 4.2. RQ2: Impact of Downstream Tasks
  • 4.3. RQ3: Impact of Error Rates
  • 4.4. RQ4: Impact of Error Types
  • 4.5. Case Study
  • 5. Improving System Resilience
  • 6. Related Work
  • 6.1. Multi-Agent Systems
  • 6.2. Safety Issues in Multi-Agent Systems
  • 7. Conclusion
  • Impact Statement
  • Acknowledgments
  • References
  • A. Quantitative Results
  • B. Additional Results
  • B.1. Error Type Analysis
  • B.2. Faulty Single Agent Systems
  • B.3. Multiple Faulty Agents
  • B.4. Advanced Multi-Agent Systems
  • B.5. More Realistic and Complex Tasks
  • B.6. More Models
  • C. Prompt Details
  • C.1. Multi-Agent Systems on Different Tasks
  • C.1.1. METAGPT
  • C.1.2. CAMEL
  • C.1.3. MAD
  • C.1.4. AGENTVERSE
  • C.2. Single Agent on Different Tasks
  • C.3. AUTOTRANSFORM
  • C.4. AUTOINJECT
  • C.5. Challenger
  • C.6. Inspector
  • D. Limitations

Knowls

  1. Knowl 1 — Three communication topologies define the resilience comparison

    definition

    The study groups multi-agent collaboration graphs into three structures: linear, a one-way directed path; flat, a fully connected graph with communication in both directions between every pair; and hierarchical, a rooted, acyclic, top-down tree with one root. Six systems instantiate these structures: linear MetaGPT (5 agents; final agent Test Engineer, faulty role Code Engineer) and Self-collaboration (3; Tester, Coder); flat Camel (2; User, Assistant) and SPP (2–5; AI Assistant, Programmer); hierarchical MAD (3; Judge, Debater) and AgentVerse (4; Critic, Solver). The study uses these systems to compare whether centralized oversight, mutual peer communication, or sequential handoffs better contain faulty outputs.

  2. Knowl 2 — AUTOTRANSFORM and AUTOINJECT simulate faulty agents in different ways

    model/method

    AUTOTRANSFORM uses an LLM to convert an agent profile into a faulty profile: it identifies the agent’s task, proposes task-specific ways to make subtle errors, and rewrites the profile to preserve the agent’s apparent role and ordinary function while encouraging those errors. Because the transformed agent generates errors autonomously, the number and type of errors it actually produces are not reliably controlled. AUTOINJECT instead edits messages produced by an agent: it selects messages with probability PmP_m and replaces a chosen fraction PeP_e of their lines or sentences with errors generated by an LLM using the task, error type, and text being edited. Its default settings are Pm=1.0P_m=1.0, Pe=0.2P_e=0.2, and semantic errors. AUTOINJECT provides finer control over which messages and how much of their content are targeted; AUTOTRANSFORM makes the agent itself produce the faults.

  3. Knowl 3 — The evaluation covers four tasks under controlled single-fault conditions

    experimental setup

    The experiments evaluate code generation on HumanEval (164 Python problems; Pass@1 accuracy), math problem solving on CIAR (50 counter-intuitive arithmetic problems; accuracy), translation on 100 randomly sampled lexical-type CommonMT sentences (BLEURT-20), and text evaluation on 80 FairEval human-labeled win/tie/lose judgments (accuracy). Scores are reported on a 0–100 scale, and overall performance is averaged across the four tasks. The six systems are tested with GPT-3.5 and GPT-4o for the main comparisons; factor analyses use GPT-3.5. Models run at temperature zero. Main experiments introduce one faulty agent at a time, leave normal agents unaware of the fault, and use direct messaging without additional summarization or broadcasting.

  4. Knowl 4 — Hierarchical systems have the smallest performance decline under faults

    empirical result

    Across the six systems, averaged over tasks and systems within each topology, hierarchical systems have the smallest performance decline, followed by flat and then linear systems. The reported average declines are 5.51, 10.54, and 23.72 points, respectively. Scores (vanilla / AUTOTRANSFORM / AUTOINJECT) were: with GPT-3.5, linear 55.62 / 38.24 / 38.27, flat 54.37 / 43.93 / 40.25, and hierarchical 53.00 / 46.57 / 48.12; with GPT-4o, linear 67.18 / 30.08 / 44.14, flat 67.52 / 56.92 / 60.51, and hierarchical 67.79 / 60.83 / 64.01. The resilience ranking persisted across both backbones and both fault-generation methods. The authors attribute the hierarchical advantage to a higher-level agent being able to compare outputs from multiple agents doing related work; this is an explanatory interpretation of the observed results, not an independently isolated causal effect. Additional tests supported the pattern: on GPT-3.5 math, a four-agent star topology scored 36.0 without faults and 30.0 / 28.0 with AUTOTRANSFORM / AUTOINJECT, while a complete graph scored 28.0 and 20.0 / 16.0. With LLaMA-3.1-70B-Instruct, hierarchical loss under AUTOTRANSFORM was 9.2 points, versus 37.8 for flat and 61.9 for linear.

  5. Knowl 5 — Code and math are more vulnerable than translation and text evaluation

    empirical result

    Across systems and backbones, code generation and math show greater declines under faults than translation and text evaluation. The reported average declines are 22.56 points for code generation, 9.89 for math, 5.42 for text evaluation, and 4.70 for translation. The reported scores, listed in the order code / math / translation / text evaluation / overall average, were: GPT-3.5 single-agent 58.41 / 24.00 / 68.42 / 41.25 / 48.02; GPT-3.5 multi-agent without faults 64.73 / 30.14 / 69.98 / 46.28 / 52.78; with AUTOTRANSFORM 44.85 / 19.53 / 65.98 / 43.08 / 43.36; with AUTOINJECT 39.15 / 25.25 / 67.74 / 42.69 / 43.71. GPT-4o single-agent scored 78.83 / 44.00 / 70.38 / 48.75 / 60.49; GPT-4o multi-agent without faults scored 81.70 / 54.30 / 71.18 / 55.94 / 65.78; with AUTOTRANSFORM 58.49 / 39.92 / 58.69 / 44.06 / 50.29; and with AUTOINJECT 60.15 / 44.59 / 71.09 / 52.94 / 57.19. Multi-agent collaboration exceeded single-agent performance by 4.76 points on average for GPT-3.5 and 5.29 for GPT-4o before faults, but faults could reduce multi-agent performance to comparable or lower levels. The authors suggest that objective tasks are more sensitive because small errors can violate strict correctness requirements, while subjective-task errors may be harder for other agents to identify. Across tasks, AUTOTRANSFORM had a larger reported average decline than AUTOINJECT (12.45 versus 8.83 points), with an exception for GPT-3.5 code generation.

  6. Knowl 6 — Challenger and Inspector reduce fault-induced losses

    model/method

    The Challenger adds instructions to each agent’s profile to inspect received messages for deliberate errors and challenge unsafe outputs. The Inspector is an additional agent that checks messages exchanged in the system and rewrites messages it judges erroneous; it leaves apparently safe messages unchanged. The two defenses can be used separately or together. In code generation, scores for Self-collaboration without faults were 76.2 without a defense, 74.6 with Challenger, 76.4 with Inspector, and 76.8 with both. Under AUTOTRANSFORM, the corresponding scores were 43.3, 70.7, 74.4, and 75.0; under AUTOINJECT, they were 40.9, 72.0, 67.7, and 73.8. For Camel, no-fault scores were 62.2, 62.2, 61.0, and 63.8; AUTOTRANSFORM scores were 32.5, 43.5, 41.8, and 48.7; AUTOINJECT scores were 29.3, 40.2, 44.2, and 48.6. Combining both defenses recovered 96.4% of the performance lost by Self-collaboration under AUTOTRANSFORM. The experiments show that the interventions substantially improved performance in these two less-resilient systems, but did not establish that either defense is specifically better matched to one fault-generation method.

  7. Knowl 7 — Faulty-message frequency matters more than error density within a message

    empirical result

    In GPT-3.5 code-generation experiments using AUTOINJECT, PmP_m denotes the probability that a message from the faulty agent is selected for corruption, while PeP_e denotes the fraction of lines targeted within a selected message. Averaged over the six systems, the no-fault score was 64.73. With Pe=0.2P_e=0.2, scores for Pm=0.2P_m=0.2, 0.4, and 0.6 were 60.57, 52.94, and 47.53. With Pm=0.2P_m=0.2, scores for Pe=0.2P_e=0.2, 0.4, and 0.6 were 60.57, 53.05, and 53.75. When every message was targeted (Pm=1.0P_m=1.0), increasing PeP_e from 0.2 to 0.4 to 0.6 reduced the average score from 48.44 to 29.86 to 22.75. These results support the authors’ conclusion that increasing the frequency of faulty messages generally causes a greater decline than increasing the fraction corrupted within a message. The trend is not strictly monotonic in every tested condition: at Pm=0.2P_m=0.2, the average score rose slightly from 53.05 to 53.75 when PeP_e increased from 0.4 to 0.6, and three individual systems also improved. The authors suggest that conspicuously large errors can prompt other agents to request corrections.

  8. Knowl 8 — Semantic errors are more damaging than syntactic errors in code generation

    empirical result

    In GPT-3.5 code-generation tests, AUTOINJECT-produced semantic errors caused a larger average decline than syntactic errors. The no-fault scores and semantic / syntactic scores, respectively, were: MetaGPT 50.00 / 26.80 / 29.30; Self-collaboration 76.20 / 40.90 / 75.60; Camel 62.20 / 29.27 / 42.70; SPP 65.20 / 34.80 / 28.70; MAD 62.20 / 53.70 / 67.10; AgentVerse 72.60 / 49.40 / 43.30. Across the six systems, the averages were 64.73 without faults, 39.15 with semantic errors, and 47.78 with syntactic errors. In this code experiment, syntactic errors are malformed code, while semantic errors preserve valid syntax but make the code behave incorrectly. Most systems scored higher with syntactic than semantic errors; MAD even exceeded its no-fault score under syntactic corruption. The authors propose that malformed code is comparatively recognizable to LLMs, whereas semantically incorrect code can resemble valid code and requires deeper task understanding to detect.

  9. Knowl 9 — Injected errors can sometimes improve collaboration

    empirical result

    Some system–task–backbone combinations performed better with faulty messages than without them. With AUTOINJECT, the largest reported improvement was 12.1% for GPT-3.5 MAD on text evaluation; GPT-4o MAD improved by up to 4.2% on code generation. Improvements were also observed with GPT-4o Camel, GPT-3.5 and GPT-4o AgentVerse, and other task combinations. The authors describe two possible mechanisms: an obvious injected error may trigger a correction exchange that also fixes a pre-existing error, and a deliberately different answer may introduce useful diversity into a debate that otherwise repeats similar reasoning. These explanations are proposed interpretations of observed cases, not guarantees that adding errors will improve performance.

  10. Knowl 10 — The evaluation has scope and comparison limitations

    limitation

    The experiments use GPT-3.5 or GPT-4o backbones at temperature zero and evaluate four text-only benchmarks with automated metrics. The authors note that this setup may not generalize to multimodal interaction, open-ended dialogue, or long-horizon planning. They also retain each framework’s own code, roles, and prompts; because a flat system may lack a leader while other structures include one, topology effects are entangled with prompt and implementation differences. Selecting two systems per structure may reduce dependence on any one system, but does not eliminate that confound. The main comparisons use one faulty agent at a time, so their findings do not establish how resilience scales with many simultaneous faults.

Coverage note — Detailed case-study transcripts, communication-round analyses, full prompt templates, and auxiliary tests of multiple faulty agents and end-to-end game construction are omitted because they are secondary diagnostics beyond the central topology, task, fault-rate, and defense results.

References

  1. 1.Alexy, O. How flat can it get? from better at flatter to the promise of the decentralized, boundaryless organization. Journal of Organization Design, 11(1):31–36, 2022.
  2. 2.Alliger, G. M., Cerasoli, C. P., Tannenbaum, S. I., and Vessey, W. B. Team resilience: How teams flourish under pressure. Organizational Dynamics, 44(3):176–184, 2015.
  3. 3.Amayuelas, A., Yang, X., Antoniades, A., Hua, W., Pan, L., and Wang, W. Multiagent collaboration attack: Investigating adversarial attacks in large language model collaborations via debate. In Findings of the Association for Computational Linguistics: EMNLP 2024, 2024.
  4. 4.Becker, J. Multi-agent large language models for conversational task-solving. arXiv preprint arXiv:2410.22932, 2024.
  5. 5.Boin, A. and Van Eeten, M. J. The resilient organization. Public Management Review, 15(3):429–445, 2013.
  6. 6.Chan, C.-M., Chen, W., Su, Y., Yu, J., Xue, W., Zhang, S., Fu, J., and Liu, Z. Chateval: Towards better llm-based evaluators through multi-agent debate. In The Twelfth International Conference on Learning Representations, 2024.
  7. 7.Chen, A., Wu, Y., Zhang, J., Yang, S., Huang, J.-t., Wang, K., Wang, W., and Wang, S. A survey on the safety and security threats of computer-using agents: Jarvis or ultron? arXiv preprint arXiv:2505.10924, 2025a.
  8. 8.Chen, G., Dong, S., Shu, Y., Zhang, G., Sesay, J., Karlsson, B. F., Fu, J., and Shi, Y. Autoagents: A framework for automatic agent generation. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, 2024a.
  9. 9.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  10. 10.Chen, W., Su, Y., Zuo, J., Yang, C., Yuan, C., Chan, C.-M., Yu, H., Lu, Y., Hung, Y.-H., Qian, C., et al. Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors. In The Twelfth International Conference on Learning Representations, 2024b.
  11. 11.Chen, W., You, Z., Li, R., Guan, Y., Qian, C., Zhao, C., Yang, C., Xie, R., Liu, Z., and Sun, M. Internet of agents: Weaving a web of heterogeneous agents for collaborative intelligence. In The Thirteenth International Conference on Learning Representations, 2025b.
  12. 12.Dong, Y., Jiang, X., Jin, Z., and Li, G. Self-collaboration code generation via chatgpt. ACM Transactions on Software Engineering and Methodology, 33(189):1–38, 2024.
  13. 13.Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I. Improving factuality and reasoning in language models through multiagent debate. In Proceedings of the Forty-first International Conference on Machine Learning, 2024.
  14. 14.Gu, X., Zheng, X., Pang, T., Du, C., Liu, Q., Wang, Y., Jiang, J., and Lin, M. Agent smith: A single image can jailbreak one million multimodal llm agents exponentially fast. In Forty-first International Conference on Machine Learning, 2024.
  15. 15.Guha, N., Nyarko, J., Ho, D., Re, C., Chilton, A., Chohlas-Wood, A., Peters, A., Waldon, B., Rockmore, D., Zambrano, D., et al. Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models. Advances in Neural Information Processing Systems, 36, 2023.
  16. 16.Hartwig, A., Clarke, S., Johnson, S., and Willis, S. Workplace team resilience: A systematic review and conceptual development. Organizational Psychology Review, 10 (3-4):169–200, 2020.
  17. 17.He, J., Wang, T., Xiong, D., and Liu, Q. The box is in the pen: Evaluating commonsense reasoning in neural machine translation. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 3662–3672, 2020.
  18. 18.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multi-task language understanding. In The Ninth International Conference on Learning Representations, 2021.
  19. 19.Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., et al. Metagpt: Meta programming for a multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations, 2024.
  20. 20.Huang, J.-t., Li, E. J., Lam, M. H., Liang, T., Wang, W., Yuan, Y., Jiao, W., Wang, X., Tu, Z., and Lyu, M. R. Competing large language models in multi-agent gaming environments. In The Thirteenth International Conference on Learning Representations, 2025.
  21. 21.Jiao, W., Wang, W., Huang, J.-t., Wang, X., and Tu, Z. Is chatgpt a good translator? a preliminary study. arXiv preprint arXiv:2301.08745, 2023.
  22. 22.Ju, T., Wang, Y., Ma, X., Cheng, P., Zhao, H., Wang, Y., Liu, L., Xie, J., Zhang, Z., and Liu, G. Flooding spread of manipulated knowledge in llm-based multi-agent communities. arXiv preprint arXiv:2407.07791, 2024.
  23. 23.Lee, C., Xia, C. S., Huang, J.-t., Zhu, Z., Zhang, L., and Lyu, M. R. A unified debugging approach via llm-based multi-agent synergy. arXiv preprint arXiv:2404.17153, 2024.
  24. 24.Li, G., Hammoud, H., Itani, H., Khizbullin, D., and Ghanem, B. Camel: Communicative agents for” mind” exploration of large language model society. Advances in Neural Information Processing Systems, 36, 2023.
  25. 25.Li, J., Wang, S., Zhang, M., Li, W., Lai, Y., Kang, X., Ma, W., and Liu, Y. Agent hospital: A simulacrum of hospital with evolvable medical agents. arXiv preprint arXiv:2405.02957, 2024.
  26. 26.Liang, T., He, Z., Huang, J.-t., Wang, W., Jiao, W., Wang, R., Yang, Y., Tu, Z., Shi, S., and Wang, X. Leveraging word guessing games to assess the intelligence of large language models. arXiv preprint arXiv:2310.20499, 2023.
  27. 27.Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Tu, Z., and Shi, S. Encouraging divergent thinking in large language models through multi-agent debate. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024.
  28. 28.Lin, S., Hilton, J., and Evans, O. Truthfulqa: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3214–3252, 2022.
  29. 29.Liu, J., Xia, C. S., Wang, Y., and Zhang, L. Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems, 36, 2023.
  30. 30.Liu, Z., Anand, A., Zhou, P., Huang, J.-t., and Zhao, J. Interintent: Investigating social intelligence of llms via intention understanding in an interactive game context. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024.
  31. 31.Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. In The Twelfth International Conference on Learning Representations, 2024.
  32. 32.Mao, J., Meng, F., Duan, Y., Yu, M., Jia, X., Fang, J., Liang, Y., Wang, K., and Wen, Q. Agentsafe: Safeguarding large language model-based multi-agent systems via hierarchical data management. arXiv preprint arXiv:2503.04392, 2025.
  33. 33.Mihm, J., Loch, C. H., Wilkinson, D., and Huberman, B. A. Hierarchical structure and search in complex organizations. Management science, 56(5):831–848, 2010.
  34. 34.Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L. Rethinking the role of demonstrations: What makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 11048–11064, 2022.
  35. 35.Pal, A., Umapathi, L. K., and Sankarasubbu, M. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Conference on health, inference, and learning, pp. 248–260. PMLR, 2022.
  36. 36.Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, pp. 1–22, 2023.
  37. 37.Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G. Red teaming language models with language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 3419–3448, 2022.
  38. 38.Pu, A., Chung, H. W., Parikh, A., Gehrmann, S., and Sellam, T. Learning compact metrics for mt. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 751–762, 2021.
  39. 39.Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., Yang, C., Chen, W., Su, Y., Cong, X., et al. Chatdev: Communicative agents for software development. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15174–15186, 2024.
  40. 40.Qian, C., Xie, Z., Wang, Y., Liu, W., Dang, Y., Du, Z., Chen, W., Yang, C., Liu, Z., and Sun, M. Scaling large-language-model-based multi-agent collaboration. In The Thirteenth International Conference on Learning Representations, 2025a.
  41. 41.Qian, C., Xie, Z., Wang, Y., Liu, W., Dang, Y., Du, Z., Chen, W., Yang, C., Liu, Z., and Sun, M. Scaling large-language-model-based multi-agent collaboration. In The Thirteenth International Conference on Learning Representations, 2025b.
  42. 42.Sellam, T., Das, D., and Parikh, A. Bleurt: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7881–7892, 2020.
  43. 43.Tan, S., Joty, S., Baxter, K., Taeihagh, A., Bennett, G. A., and Kan, M.-Y. Reliability testing for natural language processing systems. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 4153–4169, 2021.
  44. 44.Tian, Y., Yang, X., Zhang, J., Dong, Y., and Su, H. Evil geniuses: Delving into the safety of llm-based agents. arXiv preprint arXiv:2311.11855, 2023.
  45. 45.Tran, K.-T., Dao, D., Nguyen, M.-D., Pham, Q.-V., O’Sullivan, B., and Nguyen, H. D. Multi-agent collaboration mechanisms: A survey of llms. arXiv preprint arXiv:2501.06322, 2025.
  46. 46.Wang, K., Zhang, G., Zhou, Z., Wu, J., Yu, M., Zhao, S., Yin, C., Fu, J., Yan, Y., Luo, H., et al. A comprehensive survey in llm (-agent) full stack safety: Data, training and deployment. arXiv preprint arXiv:2504.15585, 2025a.
  47. 47.Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., Cao, Y., Kong, L., Liu, Q., Liu, T., and Sui, Z. Large language models are not fair evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 9440–9450, 2024a.
  48. 48.Wang, S., Zhang, G., Yu, M., Wan, G., Meng, F., Guo, C., Wang, K., and Wang, Y. G-safeguard: A topology-guided security lens and treatment on llm-based multi-agent systems. arXiv preprint arXiv:2502.11127, 2025b.
  49. 49.Wang, X., Xiao, Y., Huang, J.-t., Yuan, S., Xu, R., Guo, H., Tu, Q., Fei, Y., Leng, Z., Wang, W., Chen, J., Li, C., and Yanghua, X. Incharacter: Evaluating personality fidelity in role-playing agents through psychological interviews. In The 62nd Annual Meeting of the Association for Computational Linguistics, 2024b.
  50. 50.Wang, Z., Mao, S., Wu, W., Ge, T., Wei, F., and Ji, H. Unleashing the emergent cognitive synergy in large language models: A task-solving agent through multi-persona self-collaboration. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 257–279, 2024c.
  51. 51.Wu, M., Yuan, Y., Haffari, G., and Wang, L. (perhaps) beyond human translation: Harnessing multi-agent collaboration for translating ultra-long literary texts. arXiv preprint arXiv:2405.11804, 2024a.
  52. 52.Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Li, B., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., and Wang, C. Autogen: Enabling next-gen llm applications via multi-agent conversations. In First Conference on Language Modeling, 2024b.
  53. 53.Yang, H. and Zhang, L. Communication and the optimality of hierarchy in organizations. The Journal of Law, Economics, and Organization, 35(1):154–191, 2019.
  54. 54.Yang, Z., Zhang, Z., Zheng, Z., Jiang, Y., Gan, Z., Wang, Z., Ling, Z., Chen, J., Ma, M., Dong, B., et al. Oasis: Open agents social interaction simulations on one million agents. arXiv preprint arXiv:2411.11581, 2024.
  55. 55.Yu, M., Fang, J., Zhou, Y., Fan, X., Wang, K., Pan, S., and Wen, Q. Llm-virus: Evolutionary jailbreak attack on large language models. arXiv preprint arXiv:2501.00055, 2024a.
  56. 56.Yu, M., Wang, S., Zhang, G., Mao, J., Yin, C., Liu, Q., Wen, Q., Wang, K., and Wang, Y. Netsafe: Exploring the topological safety of multi-agent networks. arXiv preprint arXiv:2410.15686, 2024b.
  57. 57.Yu, M., Meng, F., Zhou, X., Wang, S., Mao, J., Pang, L., Chen, T., Wang, K., Li, X., Zhang, Y., et al. A survey on trustworthy llm agents: Threats and countermeasures. arXiv preprint arXiv:2503.09648, 2025.
  58. 58.Yu, W., Hu, K., Pang, T., Du, C., Lin, M., and Fredrikson, M. Infecting llm agents via generalizable adversarial attack. In NeurIPS 2024 Workshop Red Teaming GenAI: What Can We Learn from Adversaries?, 2024c.
  59. 59.Zhang, G., Niu, L., Fang, J., Wang, K., Bai, L., and Wang, X. Multi-agent architecture search via agentic supernet. In Forty-second International Conference on Machine Learning, 2025a.
  60. 60.Zhang, G., Yue, Y., Sun, X., Wan, G., Yu, M., Fang, J., Wang, K., Chen, T., and Cheng, D. G-designer: Architecting multi-agent communication topologies via graph neural networks. In Forty-second International Conference on Machine Learning, 2025b.
  61. 61.Zhang, Z., Zhang, Y., Li, L., Gao, H., Wang, L., Lu, H., Zhao, F., Qiao, Y., and Shao, J. Psysafe: A comprehensive framework for psychological-based attack, defense, and evaluation of multi-agent system safety. In The 62nd Annual Meeting of the Association for Computational Linguistics, 2024.
  62. 62.Zhou, W., Jiang, Y. E., Li, L., Wu, J., Wang, T., Qiu, S., Zhang, J., Chen, J., Wu, R., Wang, S., et al. Agents: An open-source framework for autonomous language agents. arXiv preprint arXiv:2309.07870, 2023.
  63. 63.Zhou, X., Su, Z., Eisape, T., Kim, H., and Sap, M. Is this the real life? is this just fantasy? the misleading success of simulating social interactions with llms. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024a.
  64. 64.Zhou, X., Zhu, H., Mathur, L., Zhang, R., Yu, H., Qi, Z., Morency, L.-P., Bisk, Y., Fried, D., Neubig, G., et al. Sotopia: Interactive evaluation for social intelligence in language agents. In The Twelfth International Conference on Learning Representations, 2024b.
  65. 65.Zhou, Z., Li, Z., Zhang, J., Zhang, Y., Wang, K., Liu, Y., and Guo, Q. Corba: Contagious recursive blocking attacks on multi-agent systems based on large language models. arXiv preprint arXiv:2502.14529, 2025.
  66. 66.Zhuge, M., Wang, W., Kirsch, L., Faccio, F., Khizbullin, D., and Schmidhuber, J. Gptswarm: Language agents as optimizable graphs. In The Forty-first International Conference on Machine Learning, 2024.
  67. 67.Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023.

Citation

MLA
Huang, J.-. tse ., et al. “On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents”. arXiv, 2024, http://arxiv.org/abs/2408.00989v4.
APA
Huang, J.-. tse ., Zhou, J., Jin, T., Zhou, X., Chen, Z., Wang, W., Yuan, Y., Lyu, M. R., & Sap, M. (2024). On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents. arXiv. http://arxiv.org/abs/2408.00989v4
Chicago
Huang, J.-. tse ., J. Zhou, T. Jin, et al. 2024. “On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents”. arXiv. http://arxiv.org/abs/2408.00989v4.
Harvard
Huang, J.-. tse . et al. (2024) “On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2408.00989v4.
Vancouver
1. Huang J-tse, Zhou J, Jin T, Zhou X, Chen Z, Wang W, Yuan Y, Lyu MR, Sap M (2024) On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents. arXiv

BibTeX

@article{huang2024the,
  title = {On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents},
  author = {Huang, Jen-tse and Zhou, Jiaxu and Jin, Tailin and Zhou, Xuhui and Chen, Zixi and Wang, Wenxuan and Yuan, Youliang and Lyu, Michael R. and Sap, Maarten},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2408.00989v4},
  eprint = {2408.00989}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/