Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated
To address the issue that traditional physics engines struggle to simulate complex tasks and the decline in prediction quality over long sequences, researchers proposed a training scheme for robot world models based on reinforcement learning. This method avoids reliance on real historical data and instead uses autoregressive reasoning generated by the model itself for post-training. By introducing a multi-candidate future generation and comparison protocol, combined with a multi-view visual fidelity reward function, this scheme achieved significant results on the DROID dataset. Experimental data showed that its performance outperformed the strongest baseline: LPIPS with external cameras reduced by 14%, and SSIM with wrist cameras increased by 9.1%. It won 98% of paired comparisons and achieved an 80% human preference rate in blind tests, reaching a new level of performance.