AuraTracer智迹闻
中文

EVENT DOSSIER

Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

To address the core challenge of low data acquisition efficiency in infinite-time reinforcement learning, a study published on September 7, 2026, proposed a unified framework based on the theory of large deviations. This study established the exponential decay rate of the probability of strategy selection errors as a principled measure of data acquisition efficiency and derived its variational representation using the theory of large deviations, constructing a nested optimization problem. To address the implicit nature of the problem and its difficulty in solution, the authors proposed an explicit-constrained solvable convex relaxation method and developed a lazy one-step projection subgradient algorithm to construct an adaptive data acquisition strategy. Experimental results show that the obtained reinforcement learning algorithm is nearly robustly optimal under optimality criteria (with a difference of at most a constant factor), and this framework can be extended to linear function approximation to enhance scalability.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
arXiv

SignalsSIGNALS

Keyword heat
  • arXiv1

All reports (1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective

本文提出一种基于大偏差理论的统一框架,用于解决无限时域强化学习中数据获取效率这一核心挑战。研究将策略选择错误概率的指数衰减率确立为衡量效率的原则性指标,并通过大偏差理论推导其变分表征,形成嵌套优化问题。针对该程序隐式且难以求解的特性,作者提出了带有显式约束的可解凸松弛方法,并开发了懒惰一步投影次梯度算法以构建自适应数据获取策略。实验证明,所得强化学习算法在最优性标准下近乎鲁棒最优(至多相差常数倍),且该框架可扩展至线性函数近似以提升可扩展性。