AuraTracer智迹闻
中文

EVENT DOSSIER

Reviewer Capability Governs Rejection Targeting, Not Repair Skill: Evidence from LLM Execute-Review-Revise Pipelines

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated

A study on the execution, review, and revision pipeline of large language models confirmed that the ability to review rather than the skill in fixing errors is the core factor determining the rejection of proposals. In a test with 100 Olympic math problems, using a cross-family intermediate model as reviewers and low-capability models as executors increased the final accuracy by 12 percentage points (from 52% to 64%), without any incorrect answers; in contrast, self-review of the same model resulted in a high detection rate but a low repair rate and a significant error rejection rate. The study indicated that the low damage rate of self-review stemmed from revision inertia rather than quality. When the review ability fell below a threshold, this role became ineffective, only resulting in a doubling of token costs. This study describes results under a single configuration and should be considered a controlled pilot study, not a general conclusion.

Related eventsRELATED EVENTS

All reports (1)SOURCES

A arXiv cs.CL en 2026-09-07 12:00

Reviewer Capability Governs Rejection Targeting, Not Repair Skill: Evidence from LLM Execute-Review-Revise Pipelines

研究通过跨家族中档模型担任评审员、低能力模型担任执行者的多智能体 LLM 流水线,验证了评审能力而非修复技能决定拒稿针对性。在 100 道奥林匹克数学题测试中,使用跨家族中档评审员使最终准确率提升 12 个百分点(从 52% 升至 64%,p=0.0005),且未造成任何错误答案;相比之下,同模型自审查虽检出率最高(0.85),但修复率低(15% 对 43%)且误拒率高(35% 对 2%)。研究指出,自审查的低损伤率源于修订惯性而非评审质量:18 个被错误拒答的正确答案中,执行者已合规的 3 个全部出错,其余 15 个因未被忽略而保持原样。当能力低于阈值时,评审角色失效:最弱评审员未改变任何最终答案却使 token 成本翻倍。该研究仅描述单一配置下的结果,应视为受控试点而非关于验证阶段的普遍结论。