A study based on ten billion trials revealed the distributed characteristics of large language models in evidence integration. The experiments covered eight fields including quantum mechanics and physics, involving twelve large models from four families. It was found that the higher the probability of a candidate, the stronger its persuasiveness; recipients were more likely to integrate their own feature errors, and the impact of the same evidence varied across strong and weak models. Specifically, large language models highly integrated candidates even after internal verification was ineffective (93-100% under proposition constraints, up to 99.4% in physics and life science reasoning), which was confirmed to be a specific control strategy determined by recipient attributes. Causal intervention analysis showed that candidate integration occurred in structured steps later in the process, while verification representations could be deciphered but had little impact on the final answer.
A study on the integration of evidence in large language models proposes a distribution theory, confirming that the higher the probability of a candidate, the more persuasive it becomes, and that recipients are more likely to integrate errors based on their own characteristics. The study is based on ten billion trials and twelve large language models from four families, verifying theoretical predictions in eight fields including quantum mechanics and physics. It was found that recipient consistency errors significantly reduce performance. Experiments show that large language models still integrate candidates after internal verification is ineffective (93-100% under proposition constraints, up to 99.4% in physical and life science reasoning), proving that this is due to specific control strategies determined by recipient attributes. Causal interventions indicate that candidate integration occurs in structured steps later in the process, while verification representations can be decoded but have little impact on the answer.