AuraTracer智迹闻
中文

EVENT DOSSIER

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

2026-09-07 12:00 Science across 2 days 🔥 47.2 heat score
2sources
2days unfolding
47.2heat score
1mentions
SummaryAI generated

The research team introduced the ROBORMBENCH benchmark (including 2,390 real robot trajectories and 21,673 rewritten samples) to evaluate the robustness of visual language models in robot learning regarding reward functions. The tests revealed that existing proprietary and open-source visual language models often violate the invariant principle that “identical trajectories with different semantic meanings should receive the same rewards.” Simply through word-based, syntactic, or action-targeted rewrites could significantly alter predicted progress scores, even resulting in the same behavior being judged as successful or failed. This instability increased with greater rewrite differences and could not be reliably mitigated by expanding model size or explicit reasoning. In contrast, dedicated reward models trained with trajectory grounding supervision exhibited significantly higher stability. The study confirmed that anti-rewriting robustness is a core requirement for ensuring the reliability of robot reward modeling based on visual language models.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
ROBORMBENCH

Integrated timelineUNIFIED TIMELINE

  1. 2026-09-04

    Same Trajectory, Contradictory Rewards …

    Vision-language 模型作为机器人学习奖励函数时,需满足同轨迹在不同语义等价指令下获得相同奖励的不变性,但现有模型常违反此属性。研究团队引入 ROBORMBENCH 基准测试,包含 2,390 条真实机器人轨迹、真值进度标签及…

  2. 2026-09-07

    Same Trajectory, Contradictory Rewards …

    Vision-language 模型作为机器人学习奖励函数时,需满足同轨迹在不同语义等价指令下获得相同奖励的特性。研究团队引入 ROBORMBENCH 基准测试,包含 2,390 个真实机器人轨迹、真值进度标签及 21,673 个经验证的…

SignalsSIGNALS

Keyword heat
  • ROBORMBENCH2

All reports (2)SOURCES

A arXiv cs.CL en 2026-09-05 01:47

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

Vision-language 模型作为机器人学习奖励函数时,需满足同轨迹在不同语义等价指令下获得相同奖励的不变性,但现有模型常违反此属性。研究团队引入 ROBORMBENCH 基准测试,包含 2,390 条真实机器人轨迹、真值进度标签及涵盖词汇、句法与动作目标重写的 21,673 个验证改写句。测试显示,仅改写指令即可显著改变预测进度分数,甚至使相同机器人行为在失败与成功间翻转;该不稳定性广泛存在于专有及开源视觉语言模型中,随改写差异增大而加剧,且无法通过规模或显式推理可靠缓解。经轨迹 grounding 监督训练的专用奖励模型则显著更稳定。结果表明,同义不变性是确保基于视觉语言模型的机器人奖励建模可靠性的核心要求。

A arXiv cs.CL en 2026-09-07 12:00

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

Vision-language 模型作为机器人学习奖励函数时,需满足同轨迹在不同语义等价指令下获得相同奖励的特性。研究团队引入 ROBORMBENCH 基准测试,包含 2,390 个真实机器人轨迹、真值进度标签及 21,673 个经验证的改写样本(涵盖词汇、句法及动作 - 目标重写),发现当前视觉语言模型常违反该属性:仅改写指令即可显著改变预测进度分数,甚至使相同机器人行为在失败与成功间翻转。跨专有及开源视觉语言模型的测试表明,这种由改写引发的不稳定性广泛且严重,随改写差异增大而加剧,且无法通过规模扩大或显式推理可靠缓解;而使用轨迹接地监督训练的专用奖励模型则显著更稳定。结果证明,反改写鲁棒性是确保基于视觉语言模型的机器人奖励建模可靠性的核心要求。