On September 7, 2026, researchers released TruthInsightBench, aimed at evaluating the automation capabilities of open science discovery agents. The benchmark includes 40 blind tests derived from 40 peer-reviewed studies across 10 scientific fields, providing only neutral targets and frozen data, with hidden conclusions and expected values. The fixed large-language model judge automatically scores the evidence maturity of the agents based on six dimensions and 29 empirical indicators. Tests showed that the four coding agents running on frozen base models scored between 58.4 and 60.3 points (out of 100), with no statistical significance. Although these agents can handle analysis and documentation tasks, they lack the discriminative actions necessary to establish credible claims. The study indicates that the bottleneck lies in scientific judgment rather than coding ability, and true discovery remains difficult to achieve at present. Relevant data and scoring codes have been made open source.