Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multi-agent self-play

Multi-agent self-play is a reinforcement learning training approach in which multiple autonomous agents develop and improve their skills by competing or cooperating with copies of themselves or other concurrently learning agents in a shared environment. Instead of relying on static opponents or human demonstrations, the agents dynamically generate their own training curriculum as each discovered tactic forces opposing or partnering agents to adapt with new counter-strategies. This continuous feedback loop drives the emergence of increasingly complex behaviors, sophisticated tool usage, and coordinated team strategies without requiring manually designed task progressions.

1 item

Emergent Tool Use From Multi-Agent Autocurricula

Emergent Tool Use From Multi-Agent Autocurricula

Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, Igor Mordatch

OrganizationsGoogleOpenAI

Why you should read this

Reveals that complex behaviors like tool use can emerge spontaneously from the auto-curriculum generated by simple agents competing against each other.

Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a self-supervised autocurriculum inducing multiple distinct rounds of emergent strategy, many of which require sophisticated tool use and coordination. We find clear evidence of six emergent phases in agent strategy in our environment, each of which creates a new pressure for the opposing team to adapt; for instance, agents learn to build multi-object shelters using moveable boxes which in turn leads to agents discovering that they can overcome obstacles using ramps. We further provide evidence that multi-agent competition may scale better with increasing environment complexity and leads to behavior that centers around far more human-relevant skills than other self-supervised reinforcement learning methods such as intrinsic motivation. Finally, we propose transfer and fine-tuning as a way to quantitatively evaluate targeted capabilities, and we compare hide-and-seek agents to both intrinsic motivation and random initialization baselines in a suite of domain-specific intelligence tests.

Added

2026-01-27