An Anthropic researcher just gave us a peek at self-improving AI | TechCrunch
2026-08-28 08:00Models🔥 28.9 heat score
1sources
1days unfolding
28.9heat score
2mentions
SummaryAI generated
The team led by Anthropic researcher Chen Yuehan published the paper “Automated Researchers Can Reliably Mitigate Alignment Failures,” which demonstrates early practices of self-improvement using other AI models. In 10 specific behavior alignment benchmarks, the system improved all individual performance without reducing overall performance. The Automated Alignment Researcher (AAR) led by Chen Yuehan effectively identified efficient solutions through literature search, method development, and model training. The paper indicates that the best AAR methods outperformed directions proposed by human experts in an average of six hours, with an API inference cost of about 4 dollars per hour, far lower than 150 dollars per hour for human researchers. Although the benchmarks have limitations in reflecting actual alignment goals and literature maintenance, the results prove that automated alignment makes training近期 feasible, promoting a recursive self-improvement process and potentially leading to the gradual replacement of human AI researchers.
Anthropic 研究员陈月汉(Chen Yueh-Han)发布论文《自动化研究人员可可靠地缓解对齐失败》,展示 AI 系统利用其他 AI 模型进行自我改进的早期实践。该系统在 10 个特定行为对齐基准测试中,于不降低整体性能的前提下提升了所有单项表现。由陈月汉领导的自动化对齐研究员(AAR)通过搜索文献、提出方法并训练模型,有效筛选出高效方案。论文指出,最佳 AAR 方法平均在六小时内优于人类专家提出的方向,且每小时 API 推理成本约 4 美元,远低于人类研究人员的 150 美元。尽管基准测试反映实际对齐目标及文献维护存在局限,但结果证明自动化对齐后训练近期可能具备可行性,推动递归自我改进进程,甚至可能导致人类 AI 研究人员逐渐被取代。