托卡斯蒂最佳非线性控制中最佳反馈法 (On the Optimal Feedback Law in Stochastic Optimal Nonlinear Control) - 专知论文

会员服务 ·

0

优化器 · 控制器 · 易处理的 · Performer · 相似度 ·

2021 年 10 月 25 日

On the Optimal Feedback Law in Stochastic Optimal Nonlinear Control

翻译：托卡斯蒂最佳非线性控制中最佳反馈法

Mohamed Naveed Gul Mohamed,Suman Chakravorty,Raman Goyal,Ran Wang

from arxiv, arXiv admin note: substantial text overlap with arXiv:2002.10505, arXiv:2002.09478

We consider the problem of nonlinear stochastic optimal control. This problem is thought to be fundamentally intractable owing to Bellman's infamous "curse of dimensionality". We present a result that shows that repeatedly solving an open-loop deterministic problem from the current state, similar to Model Predictive Control (MPC), results in a feedback policy that is $O(\epsilon^4)$ near to the true global stochastic optimal policy. Furthermore, empirical results show that solving the Stochastic Dynamic Programming (DP) problem is highly susceptible to noise, even when tractable, and in practice, the MPC-type feedback law offers superior performance even for stochastic systems.

翻译：我们考虑的是非线性随机最佳控制的问题。人们认为,由于Bellman的臭名昭著的“ 维度诅咒”,这个问题根本难以解决。我们提出的结果表明,反复解决当前状态的开放环的确定性问题,类似于模型预测控制(MPC ), 导致一种接近真正的全球随机最佳政策的反馈政策($O ( epsilon’4) $ ) 。此外,实证结果显示,解决斯托克动态程序(DP)问题非常容易受到噪音的影响,即便在可移动的情况下,实际上,MPC型反馈法甚至为随机系统提供了优异的性能。

0

相关内容

优化器

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

Risk Sensitive Portfolio Optimization with Regime-Switching and Default Contagion，香港理工大学应用数学系余翔助理教授，第八届全国社会媒体处理大会SMP2019

Risk Sensitive Portfolio Optimization with Regime-Switching and Default Contagion，香港理工大学应用数学系余翔助理教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

10+阅读 · 2019年10月24日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

量化金融强化学习论文集合

量化金融强化学习论文集合

专知

14+阅读 · 2019年12月18日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Reinforcement Learning: An Introduction 2018第二版 500页

Reinforcement Learning: An Introduction 2018第二版 500页

CreateAMind

14+阅读 · 2018年4月27日

Efficient Importance Sampling via Stochastic Optimal Control for Stochastic Reaction Networks

Arxiv

0+阅读 · 2021年12月20日

Playing Against Fair Adversaries in Stochastic Games with Total Rewards

Arxiv

0+阅读 · 2021年12月18日

Stability Verification in Stochastic Control Systems via Neural Network Supermartingales

Arxiv

0+阅读 · 2021年12月17日

A Generalized Minimax Q-learning Algorithm for Two-Player Zero-Sum Stochastic Games

Arxiv

0+阅读 · 2021年12月17日

Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks

Arxiv

8+阅读 · 2018年11月21日

VIP会员

文章信息

相关主题

相关VIP内容

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

Risk Sensitive Portfolio Optimization with Regime-Switching and Default Contagion，香港理工大学应用数学系余翔助理教授，第八届全国社会媒体处理大会SMP2019

Risk Sensitive Portfolio Optimization with Regime-Switching and Default Contagion，香港理工大学应用数学系余翔助理教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

10+阅读 · 2019年10月24日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《利用人工智能对军事行动进行建模》

《利用人工智能学习、优化与推演美国海军作战部队的战略布局与分散（续文）》

机器人、无人机与实时影像：应对城市爆炸威胁的三大技术方案

《指挥官意图消息中关键概念自动提取》最新47页

相关资讯

量化金融强化学习论文集合

量化金融强化学习论文集合

专知

14+阅读 · 2019年12月18日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Reinforcement Learning: An Introduction 2018第二版 500页

Reinforcement Learning: An Introduction 2018第二版 500页

CreateAMind

14+阅读 · 2018年4月27日

相关论文

Efficient Importance Sampling via Stochastic Optimal Control for Stochastic Reaction Networks

Arxiv

0+阅读 · 2021年12月20日

Playing Against Fair Adversaries in Stochastic Games with Total Rewards

Arxiv

0+阅读 · 2021年12月18日

Stability Verification in Stochastic Control Systems via Neural Network Supermartingales

Arxiv

0+阅读 · 2021年12月17日

A Generalized Minimax Q-learning Algorithm for Two-Player Zero-Sum Stochastic Games

Arxiv

0+阅读 · 2021年12月17日

Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks

Arxiv

8+阅读 · 2018年11月21日

微信扫码咨询专知VIP会员