AuraTracer智迹闻
中文

EVENT DOSSIER

When Linguistic and Internal Confidence Diverge in Large Language Models

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

A study on 30 LLMs from three model families found that language confidence and internal confidence often differ from each other. In classification tasks, there are significant differences between the two in terms of relevance, amplitude consistency, and calibration; in generation tasks, language confidence fails to effectively track uncertainty based on semantic entropy. Fine-tuned models often report higher confidence and stronger relevance, but with greater confidence gaps and poorer calibration performance. Attitudinal prompts can improve confidence without improving alignment, while score examples can preserve sorting signals if they avoid confidence value collapse. Regression analysis showed that the distribution properties of confidence scores explain most alignment patterns, and model metadata plays a smaller role after control. The study supports the view that language confidence is a “deteriorating channel” – its dispersed distribution carries useful sorting information but cannot achieve calibration; therefore, language confidence must be evaluated through multi-axis diagnostics before it is used in downstream reliability pipelines.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
arXiv

SignalsSIGNALS

Keyword heat
  • arXiv1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

When Linguistic and Internal Confidence Diverge in Large Language Models

一项针对 30 个来自三个模型家族、涵盖 8 个分类任务和 2 个生成任务的 LLM 研究,发现语言置信度与内部置信度经常发生分歧。在分类任务中,语言置信度与基于 logits 的置信度在关联、幅度一致性和校准三个维度上均存在显著差异;生成任务中,语言置信度未能有效追踪基于语义熵的不确定性。指令微调模型常报告更高置信度且关联更强,但伴随更大的置信度差距和更差的校准效果。态度提示词会提升置信度但不改善对齐,而分数范例若避免置信值坍缩则能保留排序信号。回归分析表明,置信度分数的分布属性解释了大部分对齐模式,模型元数据在控制后作用较小。这些结果支持语言置信度为“有损通道”的观点:分散的语言置信度分布虽携带有用排序信息,但无法实现校准。因此,下游可靠性管道在使用前需通过多轴诊断评估语言置信度。