ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiation
Ziyi LiuBahar SarrafzadehPei ZhouLongqi YangJieyu ZhaoAshish Sharma
Establishes ProMediate, a theory-driven evaluation benchmark and simulation testbed that measures the intervention timing, strategies, and consensus-building effectiveness of proactive AI mediator agents in multi-party negotiations.
Artificial intelligence agents powered by large language models are increasingly used to assist individuals, but real-world decision-making often depends on group collaboration. Managing multi-party discussions requires socio-cognitive skills such as tracking diverse perspectives, de-escalating emotional conflict, and proactively guiding participants past deadlocks toward shared agreements. Despite this need, research lacks systematic benchmarks to evaluate proactive artificial intelligence mediators in dynamic, multi-party environments.
The article introduces and evaluates PROMEDIATE, the first testbed and socio-cognitive evaluation framework designed to assess proactive mediator agents in complex, multi-topic, multi-party negotiations. It evaluates how effectively these agents decide when and how to intervene to facilitate consensus and resolve conversational breakdowns.
The researchers constructed a simulation environment using six complex negotiation cases from Harvard Law School's Program on Negotiation, covering domains such as healthcare and environmental policy. Simulated participants engaged in conversations across three difficulty levels corresponding to behavioral conflict modes: accommodating, avoiding, and competing. The article tested three mediator settings (no agent, a generic baseline agent, and a socially intelligent agent) across multiple model backends. Performance was measured along socio-cognitive dimensions using a novel evaluation suite tracking dynamic consensus changes, topic-level efficiency, intervention latency, post-intervention momentum, and mediator intelligence across perceptual, emotional, cognitive, and communicative factors.
The evaluation produced several key findings. First, proactive social mediation delivers substantial value in difficult, high-conflict scenarios. In the hard setting with competing participants, the socially intelligent mediator increased consensus change by 3.6 percentage points compared to the generic baseline (10.65% versus 7.01%) while responding approximately 77% faster (3.71 seconds versus 15.98 seconds). Second, scenario difficulty determines optimal mediation strategy: while hard scenarios benefited greatly from frequent, proactive interventions, easy scenarios with accommodating participants achieved higher natural consensus (22.59% consensus gain) and were disrupted by excessive intervention. Third, reasoning-oriented models proved most effective, with the reasoning-focused model achieving the highest consensus improvement (9.34%) compared to faster alternatives. Finally, factor analysis revealed that mediator process quality does not correlate linearly with immediate short-term consensus gains, as skilled mediation often introduces constructive friction by surfacing latent disagreements to build durable alignment.
These findings demonstrate that proactive mediation cannot rely on a one-size-fits-all approach or single-metric performance targets. For organizational leaders, deploying artificial intelligence to facilitate high-stakes meetings requires adaptive agents capable of assessing conflict intensity and balancing rapid intervention against organic team deliberation. Prematurely forcing superficial consensus can undermine long-term problem-solving.
Organizations developing collaborative artificial intelligence should adopt multi-dimensional socio-cognitive benchmarks rather than relying solely on speed or simple outcome metrics. Technical teams should implement adaptive intervention policies where agents dynamically modulate intervention thresholds based on real-time group friction. Before deploying these systems in live operational settings, further validation is necessary to test mediator performance with human participants, manage risks related to toxic language escalation, and mitigate potential demographic biases.
- Paper: Theory of Mind for Multi-Agent Collaboration via Large Language Models, Huao Li et al. (2023). Its analysis of agents inferring teammates’ beliefs and intentions provides a direct foundation for ProMediate’s socio-cognitive mediation and intervention decisions.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). This survey maps LLM-agent architectures and evaluation approaches, helping situate ProMediate’s proactive mediator within the broader agent-evaluation landscape.
- Paper: Understanding Social Reasoning in Language Models with Language Models, Kanishk Gandhi et al. (2023). Its controlled evaluation of belief, desire, and action inference clarifies the social-reasoning capabilities that ProMediate’s negotiation mediator must use.
- Paper: AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, Qingyun Wu et al. (2023). AutoGen’s framework for multi-agent conversations and human oversight offers useful context for ProMediate’s conversational, intervention-capable agent design.
No sufficiently relevant recommendations were found.
