• 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
  • 邀请有礼
  • 套餐&价格
  • 历史记录
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Windows 客户端微信小程序
定价
高级版会员购买积分包购买API积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验

MuDE:基于多代理分解奖励的探索。

MuDE: Multi-agent decomposed reward-based exploration.

机构信息

Electronics and Telecommunications Research Institute (ETRI), 218 Gajeong-ro, Yuseong-gu, Daejeon, 34129, South Korea.

Electronics and Telecommunications Research Institute (ETRI), 218 Gajeong-ro, Yuseong-gu, Daejeon, 34129, South Korea.

出版信息

Neural Netw. 2024 Nov;179:106565. doi: 10.1016/j.neunet.2024.106565. Epub 2024 Jul 22.

DOI:10.1016/j.neunet.2024.106565
PMID:39111159
Abstract

In cooperative multi-agent reinforcement learning, agents jointly optimize a centralized value function based on the rewards shared by all agents and learn decentralized policies through value function decomposition. Although such a learning framework is considered effective, estimating individual contribution from the rewards, which is essential for learning highly cooperative behaviors, is difficult. In addition, it becomes more challenging when reinforcement and punishment, help in increasing or decreasing the specific behaviors of agents, coexist because the processes of maximizing reinforcement and minimizing punishment can often conflict in practice. This study proposes a novel exploration scheme called multi-agent decomposed reward-based exploration (MuDE), which preferably explores the action spaces associated with positive sub-rewards based on a modified reward decomposition scheme, thus effectively exploring action spaces not reachable by existing exploration schemes. We evaluate MuDE with a challenging set of StarCraft II micromanagement and modified predator-prey tasks extended to include reinforcement and punishment. The results show that MuDE accurately estimates sub-rewards and outperforms state-of-the-art approaches in both convergence speed and win rates.

摘要

在协同多智能体强化学习中,智能体基于所有智能体共享的奖励共同优化一个集中的价值函数,并通过价值函数分解学习分散的策略。尽管这种学习框架被认为是有效的,但从奖励中估计个体贡献对于学习高度合作的行为是很困难的。此外,当强化和惩罚共存时,这变得更加具有挑战性,因为强化和惩罚有助于增加或减少智能体的特定行为,而最大化强化和最小化惩罚的过程在实践中往往会发生冲突。本研究提出了一种名为多智能体分解奖励探索(MuDE)的新探索方案,该方案基于修改后的奖励分解方案,优先探索与正子奖励相关的动作空间,从而有效地探索现有的探索方案无法到达的动作空间。我们使用具有挑战性的 StarCraft II 微观管理任务集和扩展到包含强化和惩罚的修改版捕食者-猎物任务来评估 MuDE。结果表明,MuDE 能够准确地估计子奖励,并且在收敛速度和胜率方面都优于最先进的方法。

相似文献

1
MuDE: Multi-agent decomposed reward-based exploration.MuDE:基于多代理分解奖励的探索。
Neural Netw. 2024 Nov;179:106565. doi: 10.1016/j.neunet.2024.106565. Epub 2024 Jul 22.
2
Strangeness-driven exploration in multi-agent reinforcement learning.多智能体强化学习中的奇异驱动探索。
Neural Netw. 2024 Apr;172:106149. doi: 10.1016/j.neunet.2024.106149. Epub 2024 Jan 26.
3
LJIR: Learning Joint-Action Intrinsic Reward in cooperative multi-agent reinforcement learning.LJIR:在合作多智能体强化学习中学习联合行动内在奖励
Neural Netw. 2023 Oct;167:450-459. doi: 10.1016/j.neunet.2023.08.016. Epub 2023 Aug 22.
4
Generative subgoal oriented multi-agent reinforcement learning through potential field.基于势场的面向生成子目标的多智能体强化学习。
Neural Netw. 2024 Nov;179:106552. doi: 10.1016/j.neunet.2024.106552. Epub 2024 Jul 17.
5
Credit assignment with predictive contribution measurement in multi-agent reinforcement learning.多智能体强化学习中的信用分配与预测贡献度量。
Neural Netw. 2023 Jul;164:681-690. doi: 10.1016/j.neunet.2023.05.021. Epub 2023 May 20.
6
Optimistic sequential multi-agent reinforcement learning with motivational communication.带有激励性沟通的乐观序贯多智能体强化学习。
Neural Netw. 2024 Nov;179:106547. doi: 10.1016/j.neunet.2024.106547. Epub 2024 Jul 22.
7
Multi-agent Continuous Control with Generative Flow Networks.基于生成流网络的多智能体连续控制
Neural Netw. 2024 Jun;174:106243. doi: 10.1016/j.neunet.2024.106243. Epub 2024 Mar 20.
8
Modular deep reinforcement learning from reward and punishment for robot navigation.基于奖惩的机器人导航模块化深度强化学习。
Neural Netw. 2021 Mar;135:115-126. doi: 10.1016/j.neunet.2020.12.001. Epub 2020 Dec 8.
9
Hierarchical Attention Master-Slave for heterogeneous multi-agent reinforcement learning.分层注意力主从式异构多智能体强化学习。
Neural Netw. 2023 May;162:359-368. doi: 10.1016/j.neunet.2023.02.037. Epub 2023 Mar 4.
10
Egoism, utilitarianism and egalitarianism in multi-agent reinforcement learning.多智能体强化学习中的利己主义、功利主义和平等主义。
Neural Netw. 2024 Oct;178:106544. doi: 10.1016/j.neunet.2024.106544. Epub 2024 Jul 24.