Advanced Search
    GUO Maozu, WANG Yadong, LIU Yang, SUN Huamei. RESEARCH ON Q-LEARNING ALGORITHM BASED ON METROPOLIS CRITERIONJ. Journal of Computer Research and Development, 2002, 39(6): 684-688.
    Citation: GUO Maozu, WANG Yadong, LIU Yang, SUN Huamei. RESEARCH ON Q-LEARNING ALGORITHM BASED ON METROPOLIS CRITERIONJ. Journal of Computer Research and Development, 2002, 39(6): 684-688.

    RESEARCH ON Q-LEARNING ALGORITHM BASED ON METROPOLIS CRITERION

    • The balance between exploration and exploitation is one of the key problems when action selection is performed in Q-learning. Pure exploitations will cause the agent to reach the local optimization quickly, whereas excessive explorations will degenerate the performance of the Q-learning algorithm even if they can accelerate learning process and can avoid the local optimization. In this paper, finding the optimum policy in Q-learning is described as searching optimum solution in combinatorial optimization. Then Metropolis criterion of simulated annealing algorithm is introduced in the balance between exploration and exploitation of Q-learning, and the Q-learning algorithm based on Metropolis criterion, SA-Q-learning, is correspondingly presented. Finally, tests show that SA-Q-learning converges more quickly than Q-learning, and can avoid the degeneracy in performance due to excessive explorations.
    • loading

    Catalog

      Turn off MathJax
      Article Contents

      /

      DownLoad:  Full-Size Img  PowerPoint
      Return
      Return