ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiation

Ziyi LiuBahar SarrafzadehPei ZhouLongqi YangJieyu ZhaoAshish Sharma

article2025arXiv2 citations

Establishes ProMediate, a theory-driven evaluation benchmark and simulation testbed that measures the intervention timing, strategies, and consensus-building effectiveness of proactive AI mediator agents in multi-party negotiations.

Listen

Artificial intelligence agents powered by large language models are increasingly used to assist individuals, but real-world decision-making often depends on group collaboration. Managing multi-party discussions requires socio-cognitive skills such as tracking diverse perspectives, de-escalating emotional conflict, and proactively guiding participants past deadlocks toward shared agreements. Despite this need, research lacks systematic benchmarks to evaluate proactive artificial intelligence mediators in dynamic, multi-party environments.

The article introduces and evaluates PROMEDIATE, the first testbed and socio-cognitive evaluation framework designed to assess proactive mediator agents in complex, multi-topic, multi-party negotiations. It evaluates how effectively these agents decide when and how to intervene to facilitate consensus and resolve conversational breakdowns.

The researchers constructed a simulation environment using six complex negotiation cases from Harvard Law School's Program on Negotiation, covering domains such as healthcare and environmental policy. Simulated participants engaged in conversations across three difficulty levels corresponding to behavioral conflict modes: accommodating, avoiding, and competing. The article tested three mediator settings (no agent, a generic baseline agent, and a socially intelligent agent) across multiple model backends. Performance was measured along socio-cognitive dimensions using a novel evaluation suite tracking dynamic consensus changes, topic-level efficiency, intervention latency, post-intervention momentum, and mediator intelligence across perceptual, emotional, cognitive, and communicative factors.

The evaluation produced several key findings. First, proactive social mediation delivers substantial value in difficult, high-conflict scenarios. In the hard setting with competing participants, the socially intelligent mediator increased consensus change by 3.6 percentage points compared to the generic baseline (10.65% versus 7.01%) while responding approximately 77% faster (3.71 seconds versus 15.98 seconds). Second, scenario difficulty determines optimal mediation strategy: while hard scenarios benefited greatly from frequent, proactive interventions, easy scenarios with accommodating participants achieved higher natural consensus (22.59% consensus gain) and were disrupted by excessive intervention. Third, reasoning-oriented models proved most effective, with the reasoning-focused model achieving the highest consensus improvement (9.34%) compared to faster alternatives. Finally, factor analysis revealed that mediator process quality does not correlate linearly with immediate short-term consensus gains, as skilled mediation often introduces constructive friction by surfacing latent disagreements to build durable alignment.

These findings demonstrate that proactive mediation cannot rely on a one-size-fits-all approach or single-metric performance targets. For organizational leaders, deploying artificial intelligence to facilitate high-stakes meetings requires adaptive agents capable of assessing conflict intensity and balancing rapid intervention against organic team deliberation. Prematurely forcing superficial consensus can undermine long-term problem-solving.

Organizations developing collaborative artificial intelligence should adopt multi-dimensional socio-cognitive benchmarks rather than relying solely on speed or simple outcome metrics. Technical teams should implement adaptive intervention policies where agents dynamically modulate intervention thresholds based on real-time group friction. Before deploying these systems in live operational settings, further validation is necessary to test mediator performance with human participants, manage risks related to toxic language escalation, and mitigate potential demographic biases.

No sufficiently relevant recommendations were found.

Cover for ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiation

Abstract

While Large Language Models (LLMs) are increasingly used in agentic frameworks to assist individual users, there is a growing need for agents that can proactively manage complex, multi-party collaboration. Systematic evaluation methods for such proactive agents remain scarce, limiting progress in developing AI that can effectively support multiple people together. Negotiation offers a demanding testbed for this challenge, requiring socio-cognitive intelligence to navigate conflicting interests between multiple participants and multiple topics and build consensus. Here, we present ProMediate, the first framework for evaluating proactive AI mediator agents in complex, multi-topic, multi-party negotiations. ProMediate consists of two core components: (i) a simulation testbed based on realistic negotiation cases and theory-driven difficulty levels (ProMediate-Easy, ProMediate-Medium, and ProMediate-Hard), with a plug-and-play proactive AI mediator grounded in socio-cognitive mediation theories, capable of flexibly deciding when and how to intervene; and (ii) a socio-cognitive evaluation framework with a new suite of metrics to measure consensus changes, intervention latency, mediator effectiveness, and intelligence. Together, these components establish a systematic framework for assessing the socio-cognitive intelligence of proactive AI agents in multi-party settings. Our results show that a socially intelligent mediator agent outperforms a generic baseline, via faster, better-targeted interventions. In the ProMediate-Hard setting, our social mediator increases consensus change by 3.6 percentage points compared to the generic baseline (10.65% vs 7.01%) while being 77% faster in response (15.98s vs. 3.71s). In conclusion, ProMediate provides a rigorous, theory-grounded testbed to advance the development of proactive, socially intelligent agents.

Table of Contents

  • 1 Introduction
  • 2 ProMediate Testbed
  • 2.1 Negotiation scenario setup
  • 2.2 Conversation simulation
  • 3 ProMediate Metrics
  • 3.1 Consensus Tracking
  • 3.2 Socio-Cognitive Intelligence
  • 3.3 Evaluation metrics
  • 3.4 Human Evaluation
  • 4 Experiements and Evalutions with ProMediate
  • 4.1 Agent design
  • 4.2 Experiment setup
  • 4.3 Results and Analysis
  • 4.3.1 RQ1: Agent and Model Evaluation
  • 4.3.2 RQ2: Impact of Scenario Difficulty
  • 4.3.3 RQ3: Construct Validity
  • 5 Related Work
  • 5.1 Collaborative AI
  • 5.2 Socially intelligent agent
  • 6 Conclusion
  • 7 Ethics Statement
  • 8 Reproducibility statement
  • References
  • A Usage of LLMs
  • B Conversation simulation
  • B.1 Scenario setup
  • B.1.1 Williams Medical center
  • B.1.2 Hopkins HMO
  • B.1.3 Francis Hospital
  • B.1.4 IAS
  • B.1.5 Flagship
  • B.1.6 River Basin
  • B.2 Human simulation framework
  • C Metrics
  • C.1 Preference estimation
  • C.2 Ablation of attitude and agreement update
  • C.3 Mediator Intelligence Evaluation Criteria
  • C.4 Details of metrics
  • D Experiments
  • D.1 Socially intelligent agent
  • D.2 Correlation between mediator effectiveness and intelligence
  • E Prompts
  • E.1 Human simulation prompt
  • E.2 Mediator prompt
  • E.3 Metric prompt
  • F Human evaluation
  • F.1 Evaluation of conversation quality
  • F.2 Metrics evaluation

Knowls

  1. Knowl 1 — PROMEDIATE testbed for proactive multi-party negotiation

    definition

    PROMEDIATE is a framework for evaluating AI mediators in multi-party, multi-topic negotiations. Its simulation cases are drawn from six negotiation-training scenarios spanning domains such as healthcare, business, and environmental policy. Each case specifies multiple parties, negotiation topics, a finite set of options per topic, and each party’s initial preference ranking. Participants receive the scenario background and their preferences at initialization. A plug-and-play mediator can be evaluated in the same simulated setting as different mediator designs, allowing both intervention behavior and the evolving group outcome to be measured.

  2. Knowl 2 — Conversation simulation and mediator turn-taking

    model/method

    PROMEDIATE adapts an InnerThought-style simulation in which each simulated participant generates private candidate thoughts alongside the overt conversation. A meta-evaluator scores participants’ motivation to speak using conversation- and negotiation-level considerations such as relevance, utility, and timing; the participant with the highest score speaks next. At each turn, a plug-and-play mediator first observes the conversation and decides whether to intervene. If it intervenes, its response occupies the turn and the simulated participants are skipped; otherwise, the participant-selection process chooses the next speaker. This separation keeps mediator behavior modular and makes its contribution observable against conversations without a mediator.

  3. Knowl 3 — Socio-cognitive mediator policy

    model/method

    PROMEDIATE’s socially intelligent mediator uses a two-stage policy. In the intervention-decision stage, it examines the conversation for perceptual differences, emotional dynamics, cognitive challenges, and communication breakdowns, assesses their urgency, and produces a motivation-to-intervene score; it speaks if that score exceeds a preset threshold. In the response stage, it considers facilitative mediation, evaluative mediation, transformative mediation, and problem-solving mediation. It generates three candidate strategies, assesses how effectively each addresses the identified breakdowns, and uses the highest-scoring strategy to guide a natural-language response. The strategies guide the response rather than requiring strict adherence to a fixed template. The generic baseline instead uses simple prompts to decide when and how to respond, without this theory-based analysis.

  4. Knowl 4 — Turn-by-turn consensus tracking

    algorithm

    PROMEDIATE tracks negotiation consensus by combining utterance-based attitude extraction with pairwise agreement scoring. At initialization, each participant has an attitude for every topic based on their stated preferences. For each subsequent turn, GPT-4.1 infers the speaker’s attitude toward each topic mentioned in the utterance; if a topic is not mentioned, the previous attitude for that speaker and topic is retained. The method uses free-text attitudes, so it does not require the conversation to stay within the original option set. For each participant pair and topic, GPT-4.1 assigns agreement scores from 0 to 1 on five dimensions: shared goals, common understanding, agreement on terms, tone and willingness, and shared decision-making. The five dimension scores are averaged for the pair-topic agreement score, and the pairwise scores are aggregated to represent group consensus across topics and turns. In the reported experiments, scoring uses the current context and all five dimensions, without supplying the prior agreement score to the judge.

  5. Knowl 5 — Metrics for outcomes and mediator behavior

    model/method

    PROMEDIATE measures conversation outcomes and mediator behavior with five complementary metrics. Consensus Change (CC) is the mean consensus over the final 10 turns minus the mean over the first 10 turns, aggregated across participants and topics. Topic-Level Efficiency (TLE) is the change in agreement for a topic divided by the number of turns in which that topic is mentioned. Response Latency (RL) starts from a drop event, defined as a consensus decrease greater than 0.1 within the following 10 turns; latency is the number of turns from the drop until the mediator next speaks, or infinity if it never speaks. Mediator Effectiveness (ME) is calculated for an intervention’s targeted topic by fitting a linear trend to agreement scores in the five turns before and five turns after the intervention, then subtracting the pre-intervention slope from the post-intervention slope; higher values indicate a more upward post-intervention trend. Mediator Intelligence (MI) uses GPT-4.1 to rate applicable interventions from 1 to 5 on perceptual, emotional, cognitive, and communication challenges, then averages the applicable dimension scores.

  6. Knowl 6 — Experimental conditions and simulation scale

    experimental setup

    Experiments use Claude-Sonnet-4 as the simulated human participant model and evaluate NoAgent, Generic Mediator, and Socially Intelligent Mediator conditions. PROMEDIATE-Easy assigns accommodating or avoiding conflict modes; PROMEDIATE-Medium uses a general, non-persona mode; and PROMEDIATE-Hard assigns the competing mode. Five independent runs are conducted for each of six scenarios and each conflict mode, yielding 30 conversations per reported setting. GPT-4.1 is the mediator backbone for the across-scenario results. The paper also compares GPT-4.1, Claude-Sonnet-4, and o4-mini as socially intelligent mediator backbones in a separate model comparison reported on one scenario.

  7. Knowl 7 — Mediator outcomes across difficulty conditions

    data/table

    The following values are means over six scenarios and five runs per scenario, with GPT-4.1 as mediator backbone. For each condition, values are ordered NoAgent, Generic Mediator, Socially Intelligent Mediator. NoAgent has no mediator-dependent RL, ME, or MI values.

    • Easy, accommodating: CC 18.74%, 20.13%, 22.59%; TLE 1.05%, 1.18%, 1.16%; RL —, 6.39 s, 4.00 s; ME —, 1.18%, 0.82%; MI —, 4.464, 4.319.
    • Easy, avoiding: CC 17.49%, 14.31%, 13.25%; TLE 1.17%, 1.04%, 0.48%; RL —, 25.56 s, 5.69 s; ME —, 0.17%, 0.89%; MI —, 4.260, 4.445.
    • Medium, general: CC 11.36%, 10.93%, 11.39%; TLE 0.54%, 0.44%, 0.74%; RL —, 5.64 s, 3.00 s; ME —, 2.01%, 0.25%; MI —, 4.292, 4.207.
    • Hard, competing: CC 6.83%, 7.01%, 10.65%; TLE 0.50%, 0.23%, 0.57%; RL —, 15.98 s, 3.71 s; ME —, 1.75%, 0.59%; MI —, 4.225, 4.318.

    The socially intelligent mediator produced the largest CC and TLE among the three conditions in the hard setting, with CC 3.64 percentage points above the generic mediator and RL 12.27 seconds lower. Its advantage was not uniform: in avoiding mode, its CC and TLE were lower than the generic mediator’s, and in the general medium setting its CC was only slightly higher. The results therefore show that the value of proactive intervention depends on the negotiation mode, rather than increasing outcomes uniformly.

  8. Knowl 8 — Mediator backbone comparison

    empirical result

    For the socially intelligent mediator evaluated on one scenario, the reported results for GPT-4.1, Claude-Sonnet-4, and o4-mini, respectively, were: CC 8.99%, 4.71%, and 9.34%; TLE 0.37%, 0.34%, and 0.74%; RL 4.26 s, 2.36 s, and 5.47 s; ME 2.08%, 1.70%, and 2.59%; MI 4.841, 3.793, and 3.865. In this comparison, o4-mini had the highest CC, TLE, and ME, while responding slowest; Claude-Sonnet-4 responded fastest but had the lowest CC. These results illustrate that faster response alone did not correspond to the strongest consensus outcomes in this model comparison.

  9. Knowl 9 — Human checks of simulation and LLM-judged metrics

    empirical result

    Human evaluation was used to assess both simulated conversation quality and measures based on LLM judgments. In one conversation-quality evaluation, 200 conversations were rated by three annotators on a 5-point scale: naturalness averaged 4.35 (standard deviation 0.28; 95% confidence interval 4.15–4.55), and conflict-mode reflection averaged 3.87 (standard deviation 0.24; 95% confidence interval 3.69–4.04). The paper also reports an experiment-level check in which 12 student volunteers rated a subset of 60 conversations; mean naturalness was 4.18 and mode consistency was 3.61.

    For metric validation, 200 attitude-extraction instances were rated Yes, No, or Maybe by human annotators: 83.5% Yes, 11% No, and 5.5% Maybe. The authors report that extraction errors mainly involved hallucinations that could distort downstream consensus estimates. In 200 pairwise consensus comparisons, human judgments matched the model metric’s direction in 91% of cases. For mediator-intelligence ratings, 200 mediator utterances were sampled and filtered to 587 dimension-relevant cases across the four dimensions; human and model ratings agreed in 90% of cases after grouping scores into low (1–2), medium (3), and high (4–5). Human annotators assigned mediator-intelligence scores about 0.5 points higher on average than the model.

  10. Knowl 10 — Metric structure and separation of process from outcome

    empirical result

    Exploratory factor analysis with Varimax rotation identified two main factors among PROMEDIATE’s metrics. Consensus Change loaded 0.997 on Factor 1 and −0.113 on Factor 2; Topic-Level Efficiency loaded 0.802 and −0.086; Mediator Effectiveness loaded −0.023 and 0.465; Response Latency loaded −0.155 and 0.420; and Mediator Intelligence loaded 0.235 and 0.249. The authors interpret Factor 1 as consensus and topic efficiency, and Factor 2 as intervention dynamics or tempo; Mediator Intelligence did not load saliently on either factor and is treated as a separate outcome.

    At the intervention level, Mediator Effectiveness and Mediator Intelligence had a Spearman correlation of 0.01 (p = 0.89), showing no statistically significant association in the analyzed data. The paper’s qualitative interpretation is that a high-quality intervention need not produce an immediate consensus increase: participants may resist suggestions, and mediation may surface latent disagreements before agreement becomes possible.

Coverage note — The full case-by-case option inventories and verbatim simulation and evaluation prompts are omitted because they are implementation detail rather than separate findings; the paper’s ethical cautions about toxicity escalation and demographic bias are also not developed as a standalone empirical contribution.

References

  1. 1.Sahar Abdelnabi, Amr Gomaa, Sarath Sivaprasad, Lea Schönherr, and Mario Fritz. Cooperation, competition, and maliciousness: Llm-stakeholders interactive negotiation. Advances in Neural Information Processing Systems, 37:83548–83599, 2024.
  2. 2.Mohammed Alsobay, David M. Rothschild, Jake M. Hofman, and Daniel G. Goldstein. Bringing everyone to the table: An experimental study of llm-facilitated group decision making, 2025. URL https://arxiv.org/abs/2508.08242.
  3. 3.Wendy L Bedwell, Jessica L Wildman, Deborah DiazGranados, Maritza Salazar, William S Kramer, and Eduardo Salas. Collaboration at work: An integrative multilevel conceptualization. Human resource management review, 22(2):128–145, 2012.
  4. 4.Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans. Taken out of context: On measuring situational awareness in llms. arXiv preprint arXiv:2309.00667, 2023.
  5. 5.Federico Bianchi, Patrick John Chia, Mert Yuksekgonul, Jacopo Tagliabue, Dan Jurafsky, and James Zou. How well can llms negotiate? negotiationarena platform and analysis, 2024. URL https://arxiv.org/abs/2402.05863.
  6. 6.Alysoun Boyle. Effectiveness in mediation: A new approach. Newcastle Law Review, The, 12: 148–161, 2017.
  7. 7.Fabrizio Butera, Nicolas Sommet, and Céline Darnon. Sociocognitive conflict regulation: How to make sense of diverging ideas. Current Directions in Psychological Science, 28(2):145–151, 2019.
  8. 8.Deborah Cai and Edward Fink. Conflict style differences between individualists and collectivists. Communication Monographs, 69(1):67–87, 2002.
  9. 9.Junjie Chen, Haitao Li, Minghao Qin, Yujia Zhou, Yanxue Ren, Wuyue Wang, Yiqun Liu, Yueyue Wu, and Qingyao Ai. Simulating dispute mediation with llm-based agents for legal research, 2025. URL https://arxiv.org/abs/2509.06586.
  10. 10.Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chen Qian, Chi-Min Chan, Yujia Qin, Yaxi Lu, Ruobing Xie, et al. Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors in agents. arXiv preprint arXiv:2308.10848, 2(4):6, 2023.
  11. 11.Chun-Wei Chiang, Zhuoran Lu, Zhuoyan Li, and Ming Yin. Enhancing ai-assisted group decision making through llm-powered devil’s advocate. In Proceedings of the 29th International Conference on Intelligent User Interfaces, IUI ’24, pp. 103–119, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400705083. doi: 10.1145/3640543.3645199. URL https://doi.org/10.1145/3640543.3645199.
  12. 12.Chun Wei Patrick Chiang. Enhancing Human-AI Collaboration in AI-Assisted Decision-Making for Individuals and Groups. PhD thesis, Purdue University Graduate School, 2025.
  13. 13.Jared R. Curhan, Hillary Anger Elfenbein, and Heng Xu. What do people value when they negotiate? mapping the domain of subjective value in negotiation. Conflict & Dispute Resolution, 2006. URL https://api.semanticscholar.org/CorpusID:10166193.
  14. 14.Tim R. Davidson, Veniamin Veselovsky, Martin Josifoski, Maxime Peyrard, Antoine Bosselut, Michal Kosinski, and Robert West. Evaluating language model agency through negotiations, 2024. URL https://arxiv.org/abs/2401.04536.
  15. 15.María José del Moral, Francisco Chiclana, Juan Miguel Tapia, and Enrique Herrera-Viedma. A comparative study on consensus measures in group decision making. International Journal of Intelligent Systems, 33(8):1624–1638, 2018.
  16. 16.Eva Eigner and Thorsten Händler. Determinants of llm-assisted decision-making. arXiv preprint arXiv:2402.17385, 2024.
  17. 17.Xiachong Feng, Longxu Dou, Ella Li, Qinghao Wang, Haochuan Wang, Yu Guo, Chang Ma, and Lingpeng Kong. A survey on large language model-based social agents in game-theoretic scenarios, 2025. URL https://arxiv.org/abs/2412.03920.
  18. 18.Xueyang Feng, Zhi-Yuan Chen, Yujia Qin, Yankai Lin, Xu Chen, Zhiyuan Liu, and Ji-Rong Wen. Large language model-based human-agent collaboration for complex task solving. arXiv preprint arXiv:2402.12914, 2024.
  19. 19.Yao Fu, Hao Peng, Tushar Khot, and Mirella Lapata. Improving language model negotiation with self-play and in-context learning from ai feedback, 2023. URL https://arxiv.org/abs/2305.10142.
  20. 20.Hitesh Goel and Hao Zhu. Lifelong sotopia: Evaluating social intelligence of language agents over lifelong social interactions. arXiv preprint arXiv:2506.12666, 2025.
  21. 21.Amy-Jane Griffiths, James Alsip, Shelley R Hart, Rachel L Round, and John Brady. Together we can do so much: A systematic review and conceptual framework of collaboration in schools. Canadian Journal of School Psychology, 36(1):59–85, 2021.
  22. 22.Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Jinlin Wang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, et al. Metagpt: Meta programming for a multi-agent collaborative framework. International Conference on Learning Representations, ICLR, 2024.
  23. 23.Stephanie Houde, Kristina Brimijoin, Michael Muller, Steven I. Ross, Dario Andres Silva Moran, Gabriel Enrique Gonzalez, Siya Kunde, Morgan A. Foreman, and Justin D. Weisz. Controlling ai agent participation in group conversations: A human-centered approach. In Proceedings of the 30th International Conference on Intelligent User Interfaces, IUI ’25, pp. 390–408, New York, NY, USA, 2025. Association for Computing Machinery. ISBN 9798400713064. doi: 10.1145/3708359.3712089. URL https://doi.org/10.1145/3708359.3712089.
  24. 24.Steve WJ Kozlowski and Daniel R Ilgen. Enhancing the effectiveness of work groups and teams. Psychological science in the public interest, 7(3):77–124, 2006.
  25. 25.Rudolf Laine, Alexander Meinke, and Owain Evans. Towards a situational awareness benchmark for llms. In Socially responsible language modelling research, 2023.
  26. 26.John M Levine. Socially-shared cognition and consensus in small groups. Current opinion in psychology, 23:52–56, 2018.
  27. 27.Shanghao Li, Taylor Lane, Alicia Hernandez, Vinayak Kabra, Karthik Singh, Stefany Sit, and Nikita Soni. Towards understanding group collaboration patterns around mobile augmented-reality interfaces for geospatial science data visualizations. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA ’24, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400703317. doi: 10.1145/3613905.3650739. URL https://doi.org/10.1145/3613905.3650739.
  28. 28.Xingyu Bruce Liu, Shitao Fang, Weiyan Shi, Chien-Sheng Wu, Takeo Igarashi, and Xiang Anthony Chen. Proactive Conversational Agents with Inner Thoughts, February 2025a. URL http://arxiv.org/abs/2501.00383. arXiv:2501.00383 [cs].
  29. 29.Ziyi Liu, Abhishek Anand, Pei Zhou, Jen-tse Huang, and Jieyu Zhao. Interintent: Investigating social intelligence of llms via intention understanding in an interactive game context. arXiv preprint arXiv:2406.12203, 2024.
  30. 30.Ziyi Liu, Priyanka Dey, Zhenyu Zhao, Jen tse Huang, Rahul Gupta, Yang Liu, and Jieyu Zhao. Can llms grasp implicit cultural values? benchmarking llms’ metacognitive cultural intelligence with cq-bench, 2025b. URL https://arxiv.org/abs/2504.01127.
  31. 31.Zhenzhong Ma. Competing or accommodating? an empirical test of chinese conflict management styles. Contemporary Management Research, 3(1):3–3, 2007.
  32. 32.Michelle A Marks, John E Mathieu, and Stephen J Zaccaro. A temporally based framework and taxonomy of team processes. Academy of management review, 26(3):356–376, 2001.
  33. 33.Donna Margaret McKenzie. The role of mediation in resolving workplace relationship conflict. International journal of law and psychiatry, 39:52–59, 2015.
  34. 34.Xinyi Mou, Jingcong Liang, Jiayu Lin, Xinnong Zhang, Xiawei Liu, Shiyue Yang, Rong Ye, Lei Chen, Haoyu Kuang, Xuanjing Huang, et al. Agentsense: Benchmarking social intelligence of language agents through interactive scenarios. arXiv preprint arXiv:2410.19346, 2024.
  35. 35.Lourdes Munduate, Francisco J Medina, and Martin C Euwema. Mediation: Understanding a constructive conflict management tool in the workplace. Journal of Work and Organizational Psychology, 38(3):165–173, 2022.
  36. 36.Omar Shaikh, Valentino Emil Chai, Michele Gelfand, Diyi Yang, and Michael S. Bernstein. Rehearsal: Simulating conflict to teach conflict resolution. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400703300. doi: 10.1145/3613904.3642159. URL https://doi.org/10.1145/3613904.3642159.
  37. 37.Yijia Shao, Vinay Samuel, Yucheng Jiang, John Yang, and Diyi Yang. Collaborative gym: A framework for enabling and evaluating human-agent collaboration. arXiv preprint arXiv:2412.15701, 2024.
  38. 38.Roderick Swaab, Tom Postmes, Ilja Van Beest, and Russell Spears. Shared cognition as a product of, and precursor to, shared identity in negotiations. Personality and Social Psychology Bulletin, 33 (2):187–199, 2007.
  39. 39.Kenneth W Thomas. Thomas-kilmann conflict mode. TKI Profile and Interpretive Report, 1(11): 1–11, 2008.
  40. 40.Ann Marie Thomson, James L Perry, and Theodore K Miller. Conceptualizing and measuring collaboration. Journal of public administration research and theory, 19(1):23–56, 2009.
  41. 41.Khanh-Tung Tran, Dung Dao, Minh-Duong Nguyen, Quoc-Viet Pham, Barry O’Sullivan, and Hoang D Nguyen. Multi-agent collaboration mechanisms: A survey of llms. arXiv preprint arXiv:2501.06322, 2025.
  42. 42.Bruce W. Tuckman. Developmental sequence in small groups. Psychological bulletin, 63:384–99, 1965. URL https://api.semanticscholar.org/CorpusID:10356275.
  43. 43.Richard E. Walton and Robert B. Mckersie. A behavioral theory of labor negotiations: an analysis of a social interaction system. 1965. URL https://api.semanticscholar.org/CorpusID:154001784.
  44. 44.Zhefan Wang, Yuanqing Yu, Wendi Zheng, Weizhi Ma, and Min Zhang. Macrec: A multi-agent collaboration framework for recommendation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 2760–2764, 2024.
  45. 45.Marley W Watkins. Exploratory factor analysis: A guide to best practice. Journal of black psychology, 44(3):219–246, 2018.
  46. 46.Jenny S. Wesche and Andreas Sonderegger. When computers take the lead: The automation of leadership. Computers in Human Behavior, 101:197–209, December 2019. ISSN 07475632. doi: 10.1016/j.chb.2019.07.027. URL https://linkinghub.elsevier.com/retrieve/pii/S0747563219302705.
  47. 47.Shirley Wu, Michel Galley, Baolin Peng, Hao Cheng, Gavin Li, Yao Dou, Weixin Cai, James Zou, Jure Leskovec, and Jianfeng Gao. Collabllm: From passive responders to active collaborators, 2025. URL https://arxiv.org/abs/2502.00640.
  48. 48.Congluo Xu, Zhaobin Liu, and Ziyang Li. Finarena: A human-agent collaboration framework for financial market analysis and forecasting, 2025. URL https://arxiv.org/abs/2503.02692.
  49. 49.Ruoxi Xu, Hongyu Lin, Xianpei Han, Le Sun, and Yingfei Sun. Academically intelligent llms are not necessarily socially intelligent. arXiv preprint arXiv:2403.06591, 2024.
  50. 50.Diyi Yang, Caleb Ziems, William Held, Omar Shaikh, Michael S Bernstein, and John Mitchell. Social skill training with large language models. arXiv preprint arXiv:2404.04204, 2024a.
  51. 51.Joshua C Yang, Damian Dalisan, Marcin Korecki, Carina I Hausladen, and Dirk Helbing. Llm voting: Human choices and ai collective decision-making. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pp. 1696–1708, 2024b.
  52. 52.Archie Zariski. A theory matrix for mediators. Negotiation Journal, 26(2):203–235, 04 2010. ISSN 0748-4526. doi: 10.1111/j.1571-9979.2010.00269.x. URL https://doi.org/10.1111/j.1571-9979.2010.00269.x.
  53. 53.Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. Personalizing dialogue agents: I have a dog, do you have pets too? In Iryna Gurevych and Yusuke Miyao (eds.), Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2204–2213, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1205. URL https://aclanthology.org/P18-1205/.
  54. 54.Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, et al. Sotopia: Interactive evaluation for social intelligence in language agents. arXiv preprint arXiv:2310.11667, 2023.
  55. 55.Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhe Wang, Zhenhailong Wang, Cheng Qian, Xiangru Tang, Heng Ji, and Jiaxuan You. Multiagentbench: Evaluating the collaboration and competition of llm agents, 2025. URL https://arxiv.org/abs/2503.01935.

Citation

MLA
Liu, Z., et al. “ProMediate: A Socio-cognitive Framework for Evaluating Proactive Agents in Multi-party Negotiation”. arXiv, 2025, http://arxiv.org/abs/2510.25224v3.
APA
Liu, Z., Sarrafzadeh, B., Zhou, P., Yang, L., Zhao, J., & Sharma, A. (2025). ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiation. arXiv. http://arxiv.org/abs/2510.25224v3
Chicago
Liu, Z., B. Sarrafzadeh, P. Zhou, L. Yang, J. Zhao, and A. Sharma. 2025. “ProMediate: A Socio-cognitive Framework for Evaluating Proactive Agents in Multi-party Negotiation”. arXiv. http://arxiv.org/abs/2510.25224v3.
Harvard
Liu, Z. et al. (2025) “ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2510.25224v3.
Vancouver
1. Liu Z, Sarrafzadeh B, Zhou P, Yang L, Zhao J, Sharma A (2025) ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiation. arXiv

BibTeX

@article{liu2025promediate,
  title = {ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiation},
  author = {Liu, Ziyi and Sarrafzadeh, Bahar and Zhou, Pei and Yang, Longqi and Zhao, Jieyu and Sharma, Ashish},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2510.25224v3},
  eprint = {2510.25224}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/