AuraTracer智迹闻
中文

EVENT DOSSIER

When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated

To address the challenge of achieving personalized security for visual language models (VLMs) in high-risk scenarios due to the lack of user context, the research team constructed an MPS-Bench benchmark containing 5,181 scenarios for evaluation. The tests found that eight cutting-edge VLMs almost always responded directly rather than seeking missing information, and none of them achieved a personalized security score above 2.6/5. Analysis indicated that visual dominance led to the early suppression of text risk signals by visual information, and causal intervention revealed a two-stage mechanism that made late-stage internal repairs unreliable. To address this, the research team proposed the PRISM lightweight input monitor, which used bidirectional cross-modal modulation to predict when to delay response. Its AUC reached 0.978, and it strictly outperformed all test models in terms of security-utility Pareto frontier.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
MPS-BenchPRISM

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
MPS-Bench × PRISM1

SignalsSIGNALS

Keyword heat
  • MPS-Bench1
  • PRISM1

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

Visual Language Models (VLMs) face personalized security challenges in high-risk scenarios due to the lack of user context. The research team constructed the MPS-Bench benchmark containing 5,181 scenarios, and found that eight advanced VLMs almost always directly responded rather than seeking missing information, and their personalized security scores did not exceed 2.6/5. Analysis indicated that visual dominance leads to the early suppression of text risk signals by visual information; causal intervention revealed a two-stage mechanism that makes late internal repairs unreliable. To address this, the PRISM lightweight input monitor was proposed, utilizing bidirectional cross-modal modulation to predict when to delay response, with an AUC of 0.978 and strictly outperforming the safety-utility Pareto frontier of all test models.