Knowing When Not to Answer: Pseudo-Ensembles for Abstention in Music Audio-Language Models
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated
To address the issue of misjudgments in multiple-choice questions due to forced guessing by language models for musical audio, researchers proposed a pseudo-integration method based on a single pre-trained model to implement selective弃权. This method generates multiple prediction distributions through perturbations that do not change the correct answer, such as shuffling the order of candidate answers, corrupting audio, or swapping option labels. The average result and uncertainty metrics (such as entropy and mutual information) are calculated. In experiments on the TinyMU model on the MuChoMusic dataset, averaging the four answer orders increased the accuracy from 55.7% to 59.2%, and this uncertainty metric was more accurate in identifying errors than the single-pass entropy baseline, reducing the area under the error curve from 0.293 to 0.261. This method requires only a few additional forward passes and does not require re-training, proving the feasibility of the弃权 strategy in compact musical audio-language models.