Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
To address the core challenge of low data acquisition efficiency in infinite-time reinforcement learning, a study published on September 7, 2026, proposed a unified framework based on the theory of large deviations. This study established the exponential decay rate of the probability of strategy selection errors as a principled measure of data acquisition efficiency and derived its variational representation using the theory of large deviations, constructing a nested optimization problem. To address the implicit nature of the problem and its difficulty in solution, the authors proposed an explicit-constrained solvable convex relaxation method and developed a lazy one-step projection subgradient algorithm to construct an adaptive data acquisition strategy. Experimental results show that the obtained reinforcement learning algorithm is nearly robustly optimal under optimality criteria (with a difference of at most a constant factor), and this framework can be extended to linear function approximation to enhance scalability.