On September 7, 2026, a study published on arXiv cs.CL indicated that automatic speech recognition (ASR) systems produce fluent text that is unrelated to the content of the speech they receive, which is actually a consequence of broader grounding failures. The researchers tested two independently trained Conformer-Large recognizers (CTC and RNN-T) under conditions of environmental degradation and speaker background offset, and found that the final encoder stage is crucial: bypassing this layer leads to the divergence of almost all sentences, while bypassing the intermediate layers has little effect. At this stage, the representation becomes compact, the text can be read by the trained decoder, and element information is explicitly expressed; interference results in chaos or repetitive outputs rather than fluent falsification. The results confirmed the prerequisite for the mechanism of hallucination—the failure to generate sufficiently grounded outputs, rather than the complete origin of spontaneous hallucinations. Both studies revealed a consistent terminal stage dependence in grounded recognition across two different decoder families and various distribution offsets.
The ASR system generates fluent text (hallucinations) unrelated to the content when receiving speech. Research indicates that this is a consequence of broader grounding failures. Researchers tested two independently trained Conformer-Large recognizers (CTC and RNN-T) under conditions of environmental degradation and speaker background offset, and found that the final encoder stage is key: bypassing this layer leads to almost complete divergence of sentences, while bypassing the intermediate layers has little effect. At this stage, the representation becomes compact, the text can be read by the training decoder, and elemental information is explicitly expressed; intervention results in chaotic or repetitive outputs rather than fluent imitation. The results confirm the prerequisite for hallucinations—the failure to generate sufficiently grounded outputs, rather than the complete origin of spontaneous hallucinations. Both studies reveal a consistent end-stage dependence of grounded recognition across two different decoder families and various distribution offsets.