AuraTracer智迹闻
中文

EVENT DOSSIER

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

2026-09-07 12:00 Models across 2 days 🔥 47.2 heat score
2sources
2days unfolding
47.2heat score
1mentions
SummaryAI generated

On September 7, 2026, researchers released TruthInsightBench, aimed at evaluating the automation capabilities of open science discovery agents. The benchmark includes 40 blind tests derived from 40 peer-reviewed studies across 10 scientific fields, providing only neutral targets and frozen data, with hidden conclusions and expected values. The fixed large-language model judge automatically scores the evidence maturity of the agents based on six dimensions and 29 empirical indicators. Tests showed that the four coding agents running on frozen base models scored between 58.4 and 60.3 points (out of 100), with no statistical significance. Although these agents can handle analysis and documentation tasks, they lack the discriminative actions necessary to establish credible claims. The study indicates that the bottleneck lies in scientific judgment rather than coding ability, and true discovery remains difficult to achieve at present. Relevant data and scoring codes have been made open source.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
TruthInsightBench

Integrated timelineUNIFIED TIMELINE

  1. 2026-09-04

    TruthInsightBench: An Evidence-Grounded…

    TruthInsightBench 是一个用于自动化评估开放科学发现代理的证据 grounded 基准,包含从 10 个科学领域 40 篇同行评审研究中抽取的 40 个盲测任务。该基准仅提供中性科学目标和冻结数据,隐藏源结论、预期值和分析…

  2. 2026-09-07

    TruthInsightBench: An Evidence-Grounded…

    研究人员发布 TruthInsightBench,旨在评估开放科学发现代理的自动化能力。该基准包含 40 个盲测任务,源自 10 个科学领域的 40 篇同行评审研究,仅提供中性目标和冻结数据,隐藏结论与预期值。固定 LLM 法官依据六个维…

SignalsSIGNALS

Keyword heat
  • TruthInsightBench2

All reports (2)SOURCES

A arXiv cs.CL en 2026-09-04 20:37

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

TruthInsightBench 是一个用于自动化评估开放科学发现代理的证据 grounded 基准,包含从 10 个科学领域 40 篇同行评审研究中抽取的 40 个盲测任务。该基准仅提供中性科学目标和冻结数据,隐藏源结论、预期值和分析路径,要求代理自行确定数据支持的声明。固定 LLM 基于法官依据六个维度操作化的 29 个基于工件的项目对代理声明的证据成熟度进行评分,采用自动化确定性聚合且无人工分级。在一台冻结基础模型上,四个编码代理得分在 58.4 至 60.3 分之间(满分 100),无统计学显著差异;它们虽能胜任分析执行与文档编写,具备较强的证据可审计性和新颖性,但缺乏建立可信声明所需的区分性行为(如控制、稳健性、可证伪性及跨数据集泛化)。瓶颈在于科学判断而非编码,真实发现仍难以企及。TruthIn…

A arXiv cs.AI en 2026-09-07 12:00

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

研究人员发布 TruthInsightBench,旨在评估开放科学发现代理的自动化能力。该基准包含 40 个盲测任务,源自 10 个科学领域的 40 篇同行评审研究,仅提供中性目标和冻结数据,隐藏结论与预期值。固定 LLM 法官依据六个维度及 29 项基于实证的指标,对代理主张的证据成熟度进行自动化评分。在一款冻结基础模型上,四个编码代理得分在 58.4 至 60.3 分之间(满分 100),无统计学显著差异;它们虽能胜任分析与文档工作,但缺乏建立可信主张所需的区分性行动。研究指出瓶颈在于科学判断而非编码能力,真实发现目前难以实现。相关数据与评分代码已开源。