AuraTracer智迹闻
中文

EVENT DOSSIER

Improving Weak World Models Behind Strong Agents in Atari Pong

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
5mentions
SummaryAI generated

The research team reproduced five visual world model agents in the Atari Pong environment (DreamerV3, DIAMOND, TWISTER, Simulus, and STORM) and found significant visual or dynamic failures in their frozen world models, manifested as the disappearance of the ball, incorrect movements, and ineffective interactions between the ball and the paddle. In native zero-sample model-based reinforcement learning, new strategies trained solely using the frozen models performed significantly worse in real environments compared to the original agents: DreamerV3 showed a decline of -5.5 to -20.9, DIAMOND from 19.7 to -9.6, TWISTER from 17.7 to -13.3, Simulus from 20.8 to -11.6, and STORM from 18.7 to -21.0. Inspired by ball-related rolling failures, the researchers proposed Concept-Guided Space Regularization (CGSReg), a improved method for task-critical concept areas.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
DIAMONDDreamerV3STORMSimulusTWISTER

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
DIAMOND × DreamerV31DIAMOND × STORM1DIAMOND × Simulus1DIAMOND × TWISTER1DreamerV3 × STORM1DreamerV3 × Simulus1

SignalsSIGNALS

Keyword heat
  • DreamerV31
  • DIAMOND1
  • TWISTER1
  • Simulus1
  • STORM1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Improving Weak World Models Behind Strong Agents in Atari Pong

研究团队复现了 Atari Pong 中的五个视觉世界模型代理(DreamerV3、DIAMOND、TWISTER、Simulus 和 STORM),发现其冻结的世界模型存在明显的视觉或动力学失效,包括球消失、运动错误及无效的球 - 挡板交互。在原生零样本基于模型的强化学习(MBRL)中,仅利用冻结模型从头训练的新策略在真实环境中表现显著低于原代理:DreamerV3 下降幅度为 -5.5 至 -20.9,DIAMOND 为 19.7 至 -9.6,TWISTER 为 17.7 至 -13.3,Simulus 为 20.8 至 -11.6,STORM 为 18.7 至 -21.0(-21 为最低 Pong 回报)。受球相关滚动失败启发,研究者提出了概念引导空间正则化(CGSReg),这是一种针对任务关键概念区…