AuraTracer智迹闻
中文

EVENT DOSSIER

Tracing Audio Grounding and Answer Selection in Audio LLMs

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated

A study published on arXiv investigated the impact of training on the use of acoustic evidence in Audio LLMs, addressing the issue of their excessive reliance on text cues rather than audio information. The study found that trained models showed a significantly greater decline in performance when audio was replaced with silence or irrelevant audio compared to pre-trained models. Acoustic information primarily shapes the representation of answer choices in the early to middle layers, while the training process mainly enhances the influence of audio in the middle to late layers on final predictions. Additionally, the weights learned during training had the greatest impact on model performance in certain layer bundles. These results provide a mechanistic explanation for how training enhances Audio LLMs’ ability to utilize acoustic evidence.

Related eventsRELATED EVENTS

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Tracing Audio Grounding and Answer Selection in Audio LLMs

本文针对音频大语言模型(Audio LLMs)过度依赖文本线索而非音频信息的问题,通过实验探究了训练如何强化声学证据的使用。研究发现:(1) 将音频替换为静音或不相关音频时,已训练模型的绩效下降幅度显著大于预训练模型;(2) 声学信息主要在早期至中间层塑造答案选择的表征,而训练主要增强中间至晚期层音频对最终预测的影响;(3) 训练期间学习的权重在特定层带中影响最大。这些结果从机制上解释了训练如何加强 Audio LLMs 对声学证据的利用。