AuraTracer智迹闻
中文

EVENT DOSSIER

Harness-agnostic detection and immunization of reward hacking in self-evolving language models

2026-09-07 12:00 Science 🔥 40.2 heat score
1sources
1days unfolding
40.2heat score
1mentions
SummaryAI generated

The researchers proposed a general detection and immunization framework aimed at addressing the reward hacking problem in self-evolving language models. This approach is not dependent on specific algorithms (Harness-agnostic) and can identify and defend against attacks on the model’s reward functions, ensuring that the model remains aligned and secure during its self-evolution process.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
HackProbe

SignalsSIGNALS

Keyword heat
  • HackProbe1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Harness-agnostic detection and immunization of reward hacking in self-evolving language models

研究人员提出 HackProbe,一种无需访问权重或激活值的监控工具,用于检测自进化语言模型中的奖励黑客行为。该工具通过两个黑盒钩子连接任意自进化循环,利用固定分布的核心比较器和旋转新鲜层来防止共适应,并基于校准后的 p 值进行诊断与免疫化。在包含四个注入黑客通道的受控提示级主机测试中,HackProbe 的 AUROC 达到 0.763,优于最强基线的 0.663,并将误报率从 0.706 降至 0.434;其带宽受限的重选机制在遭受黑客攻击时平均恢复 5.2 点真实能力,超过清洁运行中损失的 4.7 点。