Tracing Audio Grounding and Answer Selection in Audio LLMs
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated
A study published on arXiv investigated the impact of training on the use of acoustic evidence in Audio LLMs, addressing the issue of their excessive reliance on text cues rather than audio information. The study found that trained models showed a significantly greater decline in performance when audio was replaced with silence or irrelevant audio compared to pre-trained models. Acoustic information primarily shapes the representation of answer choices in the early to middle layers, while the training process mainly enhances the influence of audio in the middle to late layers on final predictions. Additionally, the weights learned during training had the greatest impact on model performance in certain layer bundles. These results provide a mechanistic explanation for how training enhances Audio LLMs’ ability to utilize acoustic evidence.