AuraTracer智迹闻
中文

EVENT DOSSIER

When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated

A study on a 0.6B parameter language model found that when verifying 1,200 logical conclusions, its behavior could not distinguish between true and false (accuracy was only 50%). However, linear probes could accurately read the correct judgments in the hidden states (AUC 0.96). The study indicated that the dominant factor causing judgment loss was a single scalar: the judgment information along the alignment direction in the model’s logits was intact (AUC 0.89), but it was erased by a saturation decision threshold with an offset of +4.6 standard deviations. This diagnosis is universal; in experiments with 90 semantic label configurations and across models, behavior accuracy collapsed in a single-function relationship with the threshold offset (Spearman -0.93), while marginal rankings were less affected. The study proposed an actionable solution: by correcting with a single parameter, calibrating marginal decoding, or adjusting prompts, the threshold can be re-centered, improving the behavior accuracy of the 0.6B model from 50% to 81%, and restoring the 8B model to 94%.

Related eventsRELATED EVENTS

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models

一项针对 0.6B 语言模型的研究发现,其在验证 1,200 个逻辑结论时虽行为上无法区分真伪(准确率 50%),但线性探针可准确读取隐藏状态中的正确判决(AUC 0.96)。研究指出,导致判决丢失的主导因素是单个标量:模型输出 logits 中沿对齐方向的判决信息完整(AUC 0.89),却被一个偏移 +4.6 个标准差的饱和决策阈值擦除。该诊断具有普适性:在 90 种语义标签配置及跨模型实验中,行为准确率随阈值偏移呈单函数关系崩溃(Spearman -0.93),而边际排名受影响较小。研究提出可操作方案:通过单一参数校正、校准边际解码或调整提示词,将阈值重新中心化,可将 0.6B 模型的行为准确率从 50% 提升至 81%,使 8B 模型恢复至 94%。