Multiagent Trust Region Policy Optimization.

IEEE Trans Neural Netw Learn Syst

Published: September 2024

We extend trust region policy optimization (TRPO) to cooperative multiagent reinforcement learning (MARL) for partially observable Markov games (POMGs). We show that the policy update rule in TRPO can be equivalently transformed into a distributed consensus optimization for networked agents when the agents' observation is sufficient. By using a local convexification and trust-region method, we propose a fully decentralized MARL algorithm based on a distributed alternating direction method of multipliers (ADMM). During training, agents only share local policy ratios with neighbors via a peer-to-peer communication network. Compared with traditional centralized training methods in MARL, the proposed algorithm does not need a control center to collect global information, such as global state, collective reward, or shared policy and value network parameters. Experiments on two cooperative environments demonstrate the effectiveness of the proposed method.

Download full-text PDF

Source
http://dx.doi.org/10.1109/TNNLS.2023.3265358DOI Listing

Publication Analysis

Top Keywords

trust region
8
region policy
8
policy optimization
8
policy
5
multiagent trust
4
optimization extend
4
extend trust
4
optimization trpo
4
trpo cooperative
4
cooperative multiagent
4

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!