AuraTracer智迹闻
中文

EVENT DOSSIER

Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated

To address the issue that traditional physics engines struggle to simulate complex tasks and the decline in prediction quality over long sequences, researchers proposed a training scheme for robot world models based on reinforcement learning. This method avoids reliance on real historical data and instead uses autoregressive reasoning generated by the model itself for post-training. By introducing a multi-candidate future generation and comparison protocol, combined with a multi-view visual fidelity reward function, this scheme achieved significant results on the DROID dataset. Experimental data showed that its performance outperformed the strongest baseline: LPIPS with external cameras reduced by 14%, and SSIM with wrist cameras increased by 9.1%. It won 98% of paired comparisons and achieved an 80% human preference rate in blind tests, reaching a new level of performance.

Related eventsRELATED EVENTS

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning

研究人员提出一种基于强化学习的机器人世界模型训练方案,旨在解决传统物理引擎难以模拟复杂任务及长序列预测质量快速下降的问题。该方法通过引入强化学习后训练机制,使模型在自身生成的自回归推演上进行训练而非依赖真实历史;同时设计了多候选未来生成与比较协议、结合多视角的视觉保真度奖励函数。实验表明,该方案在 DROID 数据集上达到新的状态最水平,在所有指标上均优于最强基线(如外部相机 LPIPS 降低 14%、手腕相机 SSIM 提升 9.1%),赢得 98% 的配对比较,并在盲测中获得 80% 的人类偏好率。