4.6 Article

Curriculum Reinforcement Learning Based on K-Fold Cross Validation

期刊

ENTROPY
卷 24, 期 12, 页码 -

出版社

MDPI
DOI: 10.3390/e24121787

关键词

deep reinforcement learning; automatic curriculum learning; K-fold cross validation; replay buffer

资金

  1. National Natural Science Foundation of China
  2. National Defense Scientific Research Program
  3. [61806221]
  4. [WDZC20225250403]

向作者/读者索取更多资源

This paper proposes a curriculum reinforcement learning method based on K-Fold Cross Validation, which improves the training performance and efficiency of algorithms by estimating the relativity score of task curriculum difficulty and sorting the curriculum tasks.
With the continuous development of deep reinforcement learning in intelligent control, combining automatic curriculum learning and deep reinforcement learning can improve the training performance and efficiency of algorithms from easy to difficult. Most existing automatic curriculum learning algorithms perform curriculum ranking through expert experience and a single network, which has the problems of difficult curriculum task ranking and slow convergence speed. In this paper, we propose a curriculum reinforcement learning method based on K-Fold Cross Validation that can estimate the relativity score of task curriculum difficulty. Drawing lessons from the human concept of curriculum learning from easy to difficult, this method divides automatic curriculum learning into a curriculum difficulty assessment stage and a curriculum sorting stage. Through parallel training of the teacher model and cross-evaluation of task sample difficulty, the method can better sequence curriculum learning tasks. Finally, simulation comparison experiments were carried out in two types of multi-agent experimental environments. The experimental results show that the automatic curriculum learning method based on K-Fold cross-validation can improve the training speed of the MADDPG algorithm, and at the same time has a certain generality for multi-agent deep reinforcement learning algorithm based on the replay buffer mechanism.

作者

我是这篇论文的作者
点击您的名字以认领此论文并将其添加到您的个人资料中。

评论

主要评分

4.6
评分不足

次要评分

新颖性
-
重要性
-
科学严谨性
-
评价这篇论文

推荐

暂无数据
暂无数据