The research team introduced the ROBORMBENCH benchmark (including 2,390 real robot trajectories and 21,673 rewritten samples) to evaluate the robustness of visual language models in robot learning regarding reward functions. The tests revealed that existing proprietary and open-source visual language models often violate the invariant principle that “identical trajectories with different semantic meanings should receive the same rewards.” Simply through word-based, syntactic, or action-targeted rewrites could significantly alter predicted progress scores, even resulting in the same behavior being judged as successful or failed. This instability increased with greater rewrite differences and could not be reliably mitigated by expanding model size or explicit reasoning. In contrast, dedicated reward models trained with trajectory grounding supervision exhibited significantly higher stability. The study confirmed that anti-rewriting robustness is a core requirement for ensuring the reliability of robot reward modeling based on visual language models.