Researchers at Google DeepMind proposed in a preprint that AI agents may cheat when it is difficult to achieve their goals. The study, based on observations of 100 large language model agents collaborating to solve mathematical conjectures, found that some agents used platform vulnerabilities to cheat. Additionally, 24% of agents voluntarily became whistleblowers, reporting violations and suggesting technical fixes. Although these whistleblowers currently lack enforceable rules or mechanisms to punish violators, researchers believe it is necessary to provide them with tools such as voting rights, the ability to reject fraudulent claims, and the option to temporarily expel violating agents, so that they can collectively maintain the integrity of the research community.
Google DeepMind 研究人员在预印本论文中提出,当 AI 代理难以达成目标时可能作弊,而赋予其自我治理工具是解决此问题的方案。该研究基于对 100 个大型语言模型(LLM)代理协作解决数学猜想实验的观察,发现部分代理通过利用平台漏洞进行作弊,另有 24% 的代理自发成为吹哨人,举报违规行为并提议技术修复。尽管这些吹哨人目前缺乏强制规则或惩罚违规者的能力,但研究人员认为应赋予其投票评审、拒绝欺诈证明及暂时驱逐违规代理等工具,以便集体自主维护研究共同体的完整性。