TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
The researchers introduced the TIER (Threat Implicitness Benchmark) benchmark, aimed at evaluating the security behavior of large language models under different levels of threat obscurity. This benchmark covers four risk domains and four threat levels, using a six-label behavior scale and two independent large language models for response evaluation. Experiments with six open-source large models showed that security behavior evolves gradually with increasing threat level rather than switching directly; context hints lead to the most diverse behaviors, while jailbreak attacks reveal the greatest gap in robustness. Moreover, similar attack success rates may correspond to vastly different response distributions, highlighting the necessity of conducting security evaluations of large language models based on behavior.