When Linguistic and Internal Confidence Diverge in Large Language Models
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
A study on 30 LLMs from three model families found that language confidence and internal confidence often differ from each other. In classification tasks, there are significant differences between the two in terms of relevance, amplitude consistency, and calibration; in generation tasks, language confidence fails to effectively track uncertainty based on semantic entropy. Fine-tuned models often report higher confidence and stronger relevance, but with greater confidence gaps and poorer calibration performance. Attitudinal prompts can improve confidence without improving alignment, while score examples can preserve sorting signals if they avoid confidence value collapse. Regression analysis showed that the distribution properties of confidence scores explain most alignment patterns, and model metadata plays a smaller role after control. The study supports the view that language confidence is a “deteriorating channel” – its dispersed distribution carries useful sorting information but cannot achieve calibration; therefore, language confidence must be evaluated through multi-axis diagnostics before it is used in downstream reliability pipelines.