Policy-Iteration-Based Finite-Horizon Approximate Dynamic Programming for Continuous-Time Nonlinear Optimal Control

Abstract

The Hamilton–Jacobi–Bellman (HJB) equation serves as the necessary and sufficient condition for the optimal solution to the continuous-time (CT) optimal control problem (OCP). Compared with the infinite-horizon HJB equation, the solving of the finite-horizon (FH) HJB equation has been a long-standing challenge, because the partial time derivative of the value function is involved as an additional unknown term. To address this problem, this study first-time bridges the link between the partial time derivative and the terminal-time utility function, and thus it facilitates the use of the policy iteration (PI) technique to solve the CT FH OCPs. Based on this key finding, the FH approximate dynamic programming (ADP) algorithm is proposed leveraging an actor–critic framework. It is shown that the algorithm exhibits important properties in terms of convergence and optimality. Rather importantly, with the use of multilayer neural networks (NNs) in the actor–critic architecture, the algorithm is suitable for CT FH OCPs toward more general nonlinear and complex systems. Finally, the effectiveness of the proposed algorithm is demonstrated by conducting a series of simulations on both a linear quadratic regulator (LQR) problem and a nonlinear vehicle tracking problem.

Publication
In IEEE Transactions on Neural Networks and Learning Systems (TNNLS), 2022
{{title}}

{{snippet}}

更多内容

公开信息按发布时间滚动。阅读「亚洲威廉手机版登录」后,可回到列表或查看相邻条目。

建议先扫读标题与摘要,再进入全文。同栏目条目通常按时间倒序排列。

列表适合快速定位,正文适合核对表述。两者都保留在站内即可形成完整阅读路径。

快速通道

网站首页 · Contact · 2022 · {{title}}

正文、栏目列表与相关阅读构成完整路径,适合按主题持续查阅。