DQN–MADDPG Coordinating the Multi-agent Cooperation

Authors

  • Zhisheng Chen

DOI:

https://doi.org/10.54097/hset.v39i.6720

Keywords:

DQN; MADDPG; Multi-agent; Cooperation.

Abstract

Multi-agent coordinating aims for many agents to finish the same or more mutual goals. To achieve this function, there are two main ideas; the first is making multi-agents communicate with each other and acquire the information for other agents. Another is sharing the same environment, dividing the work, and cooperating to achieve the goal. ‘’Gym-cooking’’ is an excellent model to test the algorithms’ performance in network coordination and game theory; this is a sharing environment. Based on sharing information, the agent have two networks(policy and target) and tries to do something more efficiently. This article will increase the complexity of the environment and use different algorithms to process the experiment. Specially, this paper will use the MADDPG model as the primary model to show its performance in a complex environment and contrast other models like DQN. The MADDPG model in this experiment is unlike the traditional MADDPG; the work trains the MADDPG network to deal with emergencies and accidents.

Downloads

Download data is not yet available.

References

Rose E. Wang, Sarah A. Wu, James A. Evan, Joshua B. Tenenbaum, David C. Parkes, Max Kleiman-Weiner. Too many cooks: Bayesian inference for coordinating multi_agent collaboration, 2020,7(2):23-31.

Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, Igor Mordatch, Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments, 2020,9:323-342.

R.J.Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 1992, 8(3-4):229–256.

J. Hu and M. P. Wellman. Online learning about other agents in a dynamic multiagent system. In Proceedings of the Second International Conference on Autonomous Agents,1998, 239–246.

Craig Boutilier. Planning, learning and coordination in multiagent decision processes. In Proceedings of the 6th conference on Theoretical aspects of rationality and knowledge. Morgan Kaufmann Publishers Inc., 1996, 195–210.

A. Tampuu, T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, J. Aru, and R. Vicente. Multiagent cooperation and competition with deep reinforcement learning. PloS one, 2017, 12(4):e0172395.

T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra. Continuous control with deep reinforcement learning. 2015, arXiv preprint arXiv:1509.02971.

D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller. Deterministic policy gradient algorithms. In Proceedings of the 31st International Conference on Machine Learning, 2014, 387–395.

Eithan Ephrati and Jeffrey S Rosenschein. Divide and conquer in multi-agent planning. In AAAI, 1994. 1. 80.

Barbara J Grosz and Sarit Kraus. 1996. Collaborative plans for complex group action. Artificial Intelligence 1996, 86, 2: 269–357.

Downloads

Published

01-04-2023