AuraTracer智迹闻
中文

EVENT DOSSIER

Knowing When Not to Answer: Pseudo-Ensembles for Abstention in Music Audio-Language Models

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated

To address the issue of misjudgments in multiple-choice questions due to forced guessing by language models for musical audio, researchers proposed a pseudo-integration method based on a single pre-trained model to implement selective弃权. This method generates multiple prediction distributions through perturbations that do not change the correct answer, such as shuffling the order of candidate answers, corrupting audio, or swapping option labels. The average result and uncertainty metrics (such as entropy and mutual information) are calculated. In experiments on the TinyMU model on the MuChoMusic dataset, averaging the four answer orders increased the accuracy from 55.7% to 59.2%, and this uncertainty metric was more accurate in identifying errors than the single-pass entropy baseline, reducing the area under the error curve from 0.293 to 0.261. This method requires only a few additional forward passes and does not require re-training, proving the feasibility of the弃权 strategy in compact musical audio-language models.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
MuChoMusicTinyMU

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
MuChoMusic × TinyMU1

SignalsSIGNALS

Keyword heat
  • TinyMU1
  • MuChoMusic1

All reports (1)SOURCES

A arXiv cs.CL en 2026-09-07 12:00

Knowing When Not to Answer: Pseudo-Ensembles for Abstention in Music Audio-Language Models

针对音乐音频 - 语言模型在多项选择题中被迫猜测导致误判的问题,研究者构建了伪集成方法以实现选择性弃权。该方法利用单一预训练模型,通过打乱候选答案顺序、腐蚀音频或交换选项标签等不可改变正确答案的扰动方式生成多个预测分布,并计算其平均结果及相应的熵、互信息等不确定性度量。在 TinyMU 模型于 MuChoMusic 数据集上的实验中,对四个答案顺序取平均使准确率从 55.7% 提升至 59.2%,且该不确定性指标比单次通过熵基线更能准确识别错误,将误差保留曲线下的面积从 0.293 降至 0.261。该方法仅需额外几次前向传播且无需重新训练,使弃权策略在紧凑型音乐音频 - 语言模型中具备可行性。