Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. ACPO addresses this by introducing a novel agent-chained policy optimization approach, which effectively computes policy gradients under the Centralized Training with Decentralized Execution (CTDE) paradigm. This unlocks scalable and efficient MARL for complex tasks.
STATUS
ACTIVE
CATEGORY
Research
SOURCES
1 linked
ENTITIES
3 detected
OVERRIDE
Automated
MOMENTUM
15 days ago