EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents
2026-09-07 12:00Models🔥 47.2 heat score
2sources
1days unfolding
47.2heat score
4mentions
SummaryAI generated
The researchers proposed the EvoCUA-1.5 framework, which extends self-evolving computers’ use of agents from offline experience learning to online reinforcement learning. In this executable sandbox environment, the framework addresses the challenges of context management and sparse rewards during multiple rounds of interaction through mechanisms such as Step-Level Policy Optimization, policy-aware filtering, and dynamic adaptive curricula. Experiments show that EvoCUA-1.5 achieves a success rate of 63.2% on the OSWorld-Verified benchmark, outperforming open-source baselines of similar size and approaching the performance of large-parameter models. Additionally, related research proposes an online evolution framework that converts interaction trajectories into persistent process libraries. Results indicate that this mechanism can improve the performance of fixed computer usage stacks, but the benefits are conditional, and repeated revisions cannot guarantee restoring the original task state.