From Architecture to Output: Structural Origins of Hallucination in Large Language Models and the Amplifying Role of Data
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
A study published in September 2026 indicates that the hallucination phenomenon in large language models (i.e., generating fluent but factually incorrect outputs) is primarily caused by three architectural components: self-attention mechanism, maximum likelihood pre-training objective, and autoregressive commitment. The study proposes that attribution procedures can be performed simply by sampling access and verifies five falsifiable predictions. Direct pre-registration tests show that correct continuation with alternative replacements at divergence points can reduce downstream failure claims by 46.7 percentage points (p<10^-9), while incorrect factual replacements reduce errors at the same statistical rate. Additionally, the model only correctly answers 2.2% of alternative replacements in isolated cases. The study also finds that dataset pathology amplifies the role of each component, supporting the claim of asymmetric dependence: architectural components are necessary intermediaries for data-induced failure, but data defects are not a necessary condition for component-induced failure.