高级检索

    基于Markov对策的多Agent强化学习模型及算法研究

    RESEARCH ON MARKOV GAME-BASED MULTIAGENT REINFORCEMENT LEARNING MODEL AND ALGORITHMS

    • 摘要: 在MDP中,单Agent可以通过强化学习来寻找问题的最优解.但在多Agent系统中,MDP模型不再适用.同样极小极大Q算法只能解决采用零和对策模型的MAS学习问题.文中采用非零和Markov对策作为多Agent系统学习框架,并提出元对策强化学习的学习模型和元对策Q算法.理论证明元对策Q算法收敛在非零和Markov对策的元对策最优解.

       

      Abstract: In Markov decision process, a single agent could find the optimal policy of the problem by reinforcement learning. But the model of the MDP doesn’t adapt to the multi-agent system. And the minmax-Q learning algorithm could only solve the problem of zero-sum Markov games. In this paper, the non-zero-sum Markov games are adopted as a framework for multi-agent reinforcement learning, and the learning model and learning algorithms of the metagame reinforcement learning are brought forward. It is proved that this metagame-Q algorithms must converge at the most optimal value of the non-zero-game Markov games.

       

    /

    返回文章
    返回