智迹闻AuraTracer
EN

EVENT DOSSIER

多步骤临床LLM智能体的反事实公平性审核需要设定一个可测量的每行动不稳定性下限。

2026-09-07 12:00 科学 🔥 42.2 事件热度
1家信源
1天持续发酵
42.2事件热度
1个提及
摘要由 AI 生成

针对临床多步骤大型语言模型(LLM)代理的反事实公平性审计,必须引入经测量的单次动作不稳定性下限。仅依靠翻转率无法提供可解释性。在十六个案例场景中重复运行十次相同条件后,8.7% 的单元格中代理动作发生改变。不同动作(如 ICU 升级和受控操作等)之间的不稳定性差异高达八倍,范围从 0.022 至 0.179。

相关事件RELATED EVENTS
关键实体KEY ENTITIES
FairMedAgent

信号强度SIGNALS

关键词热度
  • FairMedAgent1

全部报道(1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor

Counterfactual audits for clinical LLM agents require a measured per-action instability floor because flip rates alone are uninterpretable. Re-running identical conditions ten times across sixteen vignettes moved agent actions in 8.7% of cells, with instability varying by a factor of eight (0.022 to 0.179) across actions like ICU escalation and controlled-s…