Flow-based Recurrent Belief State Learning for POMDPs
Partially Observab

Abstract

Partially Observable Markov Decision Process (POMDP) provides a principled and generic framework to model real world sequential decision making processes but yet remains unsolved, especially for high dimensional continuous space and unknown models. The main challenge lies in how to accurately obtain the belief state, which is the probability distribution over the unobservable environment states given historical information. Accurately calculating this belief state is a precondition for obtaining an optimal policy of POMDPs. Recent advances in deep learning techniques show great potential to learn good belief states. However, existing methods can only learn approximated distribution with limited flexibility. In this paper, we introduce the \textbf{F}l\textbf{O}w-based \textbf{R}ecurrent \textbf{BE}lief \textbf{S}tate model (FORBES), which incorporates normalizing flows into the variational inference to learn general continuous belief states for POMDPs. Furthermore, we show that the learned belief states can be plugged into downstream RL algorithms to improve performance. In experiments, we show that our methods successfully capture the complex belief states that enable multi-modal predictions as well as high quality reconstructions, and results on challenging visual-motor control tasks show that our method achieves superior performance and sample efficiency.

Publication
In International Conference on Machine Learning (ICML), 2022
{{title}}

{{snippet}}

更多内容

公开信息按发布时间滚动。阅读「亚洲威廉最新版下载」后,可回到列表或查看相邻条目。

建议先扫读标题与摘要,再进入全文。同栏目条目通常按时间倒序排列。

列表适合快速定位,正文适合核对表述。两者都保留在站内即可形成完整阅读路径。

快速通道

网站首页 · Contact · 2022 · {{title}}

正文、栏目列表与相关阅读构成完整路径,适合按主题持续查阅。