Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
For anti-factual fairness auditing of clinical multi-step large language model (LLM) agents, it is necessary to introduce a measured lower limit of single-action instability. Relying solely on the flipping rate cannot provide interpretability. After running the same conditions ten times in sixteen case scenarios, agent actions changed in 8.7% of cells. The instability difference between different actions (such as ICU upgrades and controlled operations) was as high as eight times, ranging from 0.022 to 0.179.
Counterfactual audits for clinical LLM agents require a measured per-action instability floor because flip rates alone are uninterpretable. Re-running identical conditions ten times across sixteen vignettes moved agent actions in 8.7% of cells, with instability varying by a factor of eight (0.022 to 0.179) across actions like ICU escalation and controlled-s…