AuraTracer智迹闻
中文

EVENT DOSSIER

From Answers to Interpretations: Rethinking Ambiguity-Induced Aleatoric Uncertainty Estimation in LLMs

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated

To address the problem of uncertainty estimation in existing large language models (LLMs), which relies on generating multiple clarifications and querying the model for comparative answers, resulting in redundant answers and cognitive leakage, researchers have proposed a new evaluation method. This method does not require answering the clarified inputs; instead, it directly estimates the accidental uncertainty components caused by input ambiguity from the rational interpretation space. In three benchmark tests, this new method improved the AUROC score from 60.85 to 63.34, reduced the cost of calculating output tokens by 4-26 times, decreased the number of API calls by 2.2-3.5 times, and significantly reduced the correlation between estimated values and cognitive uncertainty. The results indicate that estimating accidental uncertainty caused by ambiguity from the interpretation space is more effective than from the response space.

Related eventsRELATED EVENTS

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

From Answers to Interpretations: Rethinking Ambiguity-Induced Aleatoric Uncertainty Estimation in LLMs

本文提出一种无需生成答案即可估计大语言模型中由输入歧义引起的偶然不确定性的新方法。现有方法通过生成多种澄清并查询模型比较答案来估算该不确定性,但作者认为答案往往冗余且可能因认知泄漏产生误导。新提出的仅基于澄清的方法直接从合理解释空间估算歧义诱导分量,无需对澄清后的输入进行回答。在三个基准测试中,该方法将 AUROC 从 60.85 提升至 63.34,输出令牌计算成本降低 4-26 倍,API 调用次数降低 2.2-3.5 倍,且估计值与认知不确定性的相关性显著降低。研究结果表明,从解释空间而非响应空间估算由歧义引起的偶然不确定性更为有效。