Molecular D\'ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated
An audit of 22 state-of-the-art language models found that these models widely exhibit verbatim retrieval behavior in regression benchmark tasks. Specifically, on five specific datasets, more than 50% of the models showed such behavior; in the remaining datasets, this phenomenon occurred only in isolated cases. Experiments further indicated that changes in the reasoning level significantly altered retrieval behavior: under the same molecular and prompting conditions, models with a higher reasoning level were marked 89% more times than those with a lower level. Additionally, tests showed that the most powerful models could still recognize combinations of transformed SMILES strings and original labels in some cases. When retrieval was suppressed, the prediction errors of different models tended to be consistent, but differences in the use of verbatim reuse behavior led to varying results. These findings suggest that the general predictive ability of large language models depends not only on the number of their memory values.