Unsupervised Skill Discovery via Recurrent Skill Training

Abstract

Being able to discover diverse useful skills without external reward functions is beneficial in reinforcement learning research. Previous unsupervised skill discovery approaches mainly train different skills in parallel. Although impressive results have been provided, we found that parallel training procedure can sometimes block exploration when the state visited by different skills overlap, which leads to poor state coverage and restricts the diversity of learned skills. In this paper, we take a deeper look into this phenomenon and propose a novel framework to address this issue, which we call Recurrent Skill Training (ReST). Instead of training all the skills in parallel, ReST trains different skills one after another recurrently, along with a state coverage based intrinsic reward. We conduct experiments on a number of challenging 2D navigation environments and robotic locomotion environments. Evaluation results show that our proposed approach outperforms previous parallel training approaches in terms of state coverage and skill diversity.

Publication
In Conference on Neural Information Processing Systems (NeurIPS), 2022
{{title}}

{{snippet}}

更多内容

公开信息按发布时间滚动。阅读「亚洲威廉官网常用功能」后,可回到列表或查看相邻条目。

建议先扫读标题与摘要,再进入全文。同栏目条目通常按时间倒序排列。

列表适合快速定位,正文适合核对表述。两者都保留在站内即可形成完整阅读路径。

快速通道

网站首页 · Contact · 2022 · {{title}}

正文、栏目列表与相关阅读构成完整路径,适合按主题持续查阅。