Constraining an Unconstrained Multi-agent Policy with offline data

被引：0

作者：

Guan, Cong

Jiang, Tao

Li, Yi-Chen

Zhang, Zongzhang

Yuan, Lei

Yu, Yang ^{[1
]}

机构：

[1] Nanjing Univ, Natl Key Lab Novel Software Technol, Nanjing, Peoples R China

来源：

NEURAL NETWORKS | 2025年 / 186卷

基金：

中国国家自然科学基金;

关键词：

Multi-Agent Reinforcement Learning; Constrained reinforcement learning; Offline reinforcement learning; REINFORCEMENT; LEVEL;

D O I：

10.1016/j.neunet.2025.107253

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Real-world multi-agent decision-making systems often have to satisfy some constraints, such as harmfulness, economics, etc., spurring the emergence of Constrained Multi-Agent Reinforcement Learning (CMARL). Existing studies of CMARL mainly focus on training a constrained policy in an online manner, that is, not only maximizing cumulative rewards but also not violating constraints. However, in practice, online learning may be infeasible due to safety restrictions or a lack of high-fidelity simulators. Moreover, as the learned policy runs, new constraints, that are not taken into account during training, may occur. To deal with the above two issues, we propose a method called Constraining an UnconsTrained Multi-Agent Policy with offline data, dubbed CUTMAP, following the popular centralized training with decentralized execution paradigm. Specifically, we have formulated a scalable optimization objective within the framework of multi-agent maximum entropy reinforcement learning for CMARL. This approach is designed to estimate a decomposable Q-function by leveraging an unconstrained "prior policy"1 in conjunction with cost signals extracted from offline data. When anew constraint comes, CUTMAP can reuse the prior policy without re-training it. To tackle the distribution shift challenge in offline learning, we also incorporate a conservative loss term when updating the Q-function. Therefore, the unconstrained prior policy can be trained to satisfy cost constraints through CUTMAP without the need for expensive interactions with the real environment, facilitating the practical application of MARL algorithms. Empirical results in several cooperative multi-agent benchmarks, including StarCraft games, particle games, food search games, and robot control, demonstrate the superior performance of our method.

引用

页数：11

共 70 条

[1]

Achiam J, 2017, PR MACH LEARN RES, V70

[2]

Albrecht StefanoV., 2013, P 2013 INT C AUTONOM, P1155

[3]

Atkeson C.G., 1997, Machine Learning-International Workshop then Conference, P12

[4]

Boyd S., 2004, Convex Optimization, DOI 10.1017/CBO9780511804441

[5] A comprehensive survey of multiagent reinforcement learning [J].

Busoniu, Lucian ;

Babuska, Robert ;

De Schutter, Bart .

IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS PART C-APPLICATIONS AND REVIEWS, 2008, 38 (02) :156-172

[6]

Chen Y., 2021, arXiv

[7]

Chow Y, 2018, J MACH LEARN RES, V18

[8]

Christianos F, 2021, PR MACH LEARN RES, V139

[9] Data-driven control of hydraulic servo actuator: An event-triggered adaptive dynamic programming approach [J].

Djordjevic, Vladimir ;

Tao, Hongfeng ;

Song, Xiaona ;

He, Shuping ;

Gao, Weinan ;

Stojanovic, Vladimir .

MATHEMATICAL BIOSCIENCES AND ENGINEERING, 2023, 20 (05) :8561-8582

[10] Dynamic Event-Triggered Consensus Control for Interval Type-2 Fuzzy Multi-Agent Systems [J].

Du, Zhenbin ;

Xie, Xiangpeng ;

Qu, Zifang ;

Hu, Yangyang ;

Stojanovic, Vladimir .

IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS I-REGULAR PAPERS, 2024, 71 (08) :3857-3866

← 1 2 3 4 5 6 7 →