Reviewer Capability Governs Rejection Targeting, Not Repair Skill: Evidence from LLM Execute-Review-Revise Pipelines
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated
A study on the execution, review, and revision pipeline of large language models confirmed that the ability to review rather than the skill in fixing errors is the core factor determining the rejection of proposals. In a test with 100 Olympic math problems, using a cross-family intermediate model as reviewers and low-capability models as executors increased the final accuracy by 12 percentage points (from 52% to 64%), without any incorrect answers; in contrast, self-review of the same model resulted in a high detection rate but a low repair rate and a significant error rejection rate. The study indicated that the low damage rate of self-review stemmed from revision inertia rather than quality. When the review ability fell below a threshold, this role became ineffective, only resulting in a doubling of token costs. This study describes results under a single configuration and should be considered a controlled pilot study, not a general conclusion.