AuraTracer智迹闻
中文

EVENT DOSSIER

Molecular D\'ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated

An audit of 22 state-of-the-art language models found that these models widely exhibit verbatim retrieval behavior in regression benchmark tasks. Specifically, on five specific datasets, more than 50% of the models showed such behavior; in the remaining datasets, this phenomenon occurred only in isolated cases. Experiments further indicated that changes in the reasoning level significantly altered retrieval behavior: under the same molecular and prompting conditions, models with a higher reasoning level were marked 89% more times than those with a lower level. Additionally, tests showed that the most powerful models could still recognize combinations of transformed SMILES strings and original labels in some cases. When retrieval was suppressed, the prediction errors of different models tended to be consistent, but differences in the use of verbatim reuse behavior led to varying results. These findings suggest that the general predictive ability of large language models depends not only on the number of their memory values.

Related eventsRELATED EVENTS

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Molecular D\'ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

研究人员审计了 22 个前沿模型在 12 个回归基准上的verbatim检索情况,发现该现象广泛存在但具有特定性:在五个数据集上超过 50% 的大语言模型(LLMs)显示verbatim检索,而在其余数据集中仅出现在孤立单元格。实验表明推理水平改变检索行为,相同分子和提示下,高推理层级比最低层级多被标记 89% 的检索次数。此外,测试中断检索后发现,最强模型在部分情况下仍识别变换后的 SMILES 字符串与原始标签的组合;抑制检索使不同模型的预测误差相对趋同,而verbatim检索的使用差异则使其分散。这表明 LLM 的一般预测能力不仅取决于记忆值的数量。