Recent research indicates that agents based on large language models (LLMs) lack structural moral capabilities and therefore cannot meet the coherent policy requirements necessary for AI alignment. Researchers defined four structural criteria: decision stability, monotonicity, decisiveness, and Pareto feasibility as measures of performance. In multiple deployment tests on nine cutting-edge LLM models, it was found that none of the models exhibited consistency across scenarios: surface form perturbations alone could cause decision rates to fluctuate by up to 99 percentage points, and success in one scenario did not predict ability in other scenarios. This suggests that current LLM-based agents do not possess the prerequisites required for alignment and cannot be effectively applied with existing alignment technologies.
AI 对齐要求系统行为表达连贯政策,即映射情境与裁决且具不变性与敏感性。研究提出四项结构条件(裁决稳定性、单调性、决断力及帕累托可行性)以衡量无需道德标准即可评估的行为能力。在针对九个前沿大语言模型的五种变体、五个升级层级及三种支配条件的九次部署测试中,未发现任何模型在三类部署中表达连贯政策:表面形式扰动导致单一升级层级的裁决率波动高达 99 个百分点,且单一场景的成功无法预测另一场景的能力。这表明当前基于大语言模型的智能体尚不具备对齐所需的前提条件。