Sharp high-probability sample complexities for policy evaluation with linear function approximation - 专知论文

会员服务 ·

0

样本复杂度 · 线性的 · 策略评估 · 泛函 · 样本 ·

2023 年 5 月 30 日

Sharp high-probability sample complexities for policy evaluation with linear function approximation

翻译：暂无翻译

Gen Li,Weichen Wu,Yuejie Chi,Cong Ma,Alessandro Rinaldo,Yuting Wei

from arxiv, The first two authors contributed equally

This paper is concerned with the problem of policy evaluation with linear function approximation in discounted infinite horizon Markov decision processes. We investigate the sample complexities required to guarantee a predefined estimation error of the best linear coefficients for two widely-used policy evaluation algorithms: the temporal difference (TD) learning algorithm and the two-timescale linear TD with gradient correction (TDC) algorithm. In both the on-policy setting, where observations are generated from the target policy, and the off-policy setting, where samples are drawn from a behavior policy potentially different from the target policy, we establish the first sample complexity bound with high-probability convergence guarantee that attains the optimal dependence on the tolerance level. We also exhihit an explicit dependence on problem-related quantities, and show in the on-policy setting that our upper bound matches the minimax lower bound on crucial problem parameters, including the choice of the feature maps and the problem dimension.

翻译：暂无翻译

0

相关内容

样本复杂度

样本复杂度

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

玉米转脂蛋白新成员ZmLTP3的抗盐功能及其上游调控机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

FBXO6与糖基化Ero1L相互作用在内质网应激中的分子机制及其生物学意义

国家自然科学基金

0+阅读 · 2013年12月31日

DACT2在结直肠癌中的表达调控机制及基因功能分析

国家自然科学基金

0+阅读 · 2013年12月31日

牛磺酸对PUMA介导缺血再灌注心肌细胞凋亡的抑制作用

国家自然科学基金

0+阅读 · 2012年12月31日

地下水区间型不确定性数值模拟

国家自然科学基金

0+阅读 · 2012年12月31日

Additive Noise Mechanisms for Making Randomized Approximation Algorithms Differentially Private

Arxiv

0+阅读 · 2023年7月18日

Natural Actor-Critic for Robust Reinforcement Learning with Function Approximation

Arxiv

0+阅读 · 2023年7月17日

Asymptotic properties and approximation of Bayesian logspline density estimators for communication-free parallel computing methods

Arxiv

0+阅读 · 2023年7月16日

An Incremental Span-Program-Based Algorithm and the Fine Print of Quantum Topological Data Analysis

Arxiv

0+阅读 · 2023年7月13日

Near-Optimal Bounds for Learning Gaussian Halfspaces with Random Classification Noise

Arxiv

0+阅读 · 2023年7月13日

VIP会员

文章信息

相关主题

样本复杂度

相关VIP内容

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

热门VIP内容

开通专知VIP会员享更多权益服务

前沿人工智能趋势报告（Frontier AI Trends Report）

【AAAI2026】善始则事半功倍：基于前缀优化的大语言模型推理强化学习

Andrej Karpathy：2025 年 LLM 年度回顾（2025 LLM Year in Review）

音退化问题：基于输入操控的鲁棒语音转换综述

相关资讯

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

相关论文

Additive Noise Mechanisms for Making Randomized Approximation Algorithms Differentially Private

Arxiv

0+阅读 · 2023年7月18日

Natural Actor-Critic for Robust Reinforcement Learning with Function Approximation

Arxiv

0+阅读 · 2023年7月17日

Asymptotic properties and approximation of Bayesian logspline density estimators for communication-free parallel computing methods

Arxiv

0+阅读 · 2023年7月16日

An Incremental Span-Program-Based Algorithm and the Fine Print of Quantum Topological Data Analysis

Arxiv

0+阅读 · 2023年7月13日

Near-Optimal Bounds for Learning Gaussian Halfspaces with Random Classification Noise

Arxiv

0+阅读 · 2023年7月13日

相关基金

玉米转脂蛋白新成员ZmLTP3的抗盐功能及其上游调控机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

FBXO6与糖基化Ero1L相互作用在内质网应激中的分子机制及其生物学意义

国家自然科学基金

0+阅读 · 2013年12月31日

DACT2在结直肠癌中的表达调控机制及基因功能分析

国家自然科学基金

0+阅读 · 2013年12月31日

牛磺酸对PUMA介导缺血再灌注心肌细胞凋亡的抑制作用

国家自然科学基金

0+阅读 · 2012年12月31日

地下水区间型不确定性数值模拟

国家自然科学基金

0+阅读 · 2012年12月31日

微信扫码咨询专知VIP会员