AuraTracer智迹闻
中文

EVENT DOSSIER

KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated

In September 2026, researchers launched the KCSAT-ML benchmark. This dataset includes 664 math problems from the Korean University Entrance Examination (KCSAT) from 2014 to 2025, along with official error rates. The DRG metric was introduced to evaluate whether model errors are concentrated on problems that humans consider difficult. Experiments revealed three patterns: low-budget models experienced a collapse in accuracy at the tail of the human error-rate; scaling during testing led to a linear increase in token usage and human error rate, while accuracy improved in a non-monotonic curve; within the same model family, scaling during testing showed counter-scaling on the most difficult problems and overthinking on the simpler ones. The DRG metric indicated that models with similar accuracy might have completely opposite DRG values: one model got errors on problems humans considered difficult, while another solved difficult problems but made mistakes on those humans considered simple. The related code and dataset construction tools have been open-sourced.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
KCSAT-MLKorean College Scholastic Ability TestNaver AI

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
KCSAT-ML × Korean Colle…1KCSAT-ML × Naver AI1Korean College Scholast…1

SignalsSIGNALS

Keyword heat
  • KCSAT-ML1
  • Naver AI1
  • Korean College Scholastic Ability Test1

All reports (1)SOURCES

A arXiv cs.CL en 2026-09-07 12:00

KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty

研究人员推出 KCSAT-ML 基准测试,该数据集包含韩国大学入学能力考试(KCSAT)2014 至 2025 年共 664 道数学题及官方错误率。配套引入难度对齐推理增益(DRG)指标,用于评估模型错误是否集中在人类认为困难的题目上。实验发现三个模式:低预算模型在人类高错误率尾部准确率崩溃;测试时缩放使 token 使用与人群错误率呈线性增长,而准确率提升呈非单调曲线;同一模型家族内,测试时缩放在最难题目上表现为反缩放,在简单题目上表现为过度思考。DRG 指标显示,准确率相近的模型可能在 DRG 值上截然相反:一个模型错在人类也认为难的题,另一个则解出难题却错在人认为简单的题。相关代码与数据集构建工具已开源。