AuraTracer智迹闻
中文

EVENT DOSSIER

Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

Researchers have proposed a new type of hidden attack paradigm called “Neutral Prompt Attack” (NPA). This attack utilizes semantically benign instructions (such as encouraging imagination and detail) to redirect the generation behavior of large language models towards speculative names without specifying specific malicious package names. In evaluations against code-generation LLMs and package hallucination benchmarks, NPA significantly improved hallucination accuracy (Hallucination ASR) and pip installation accuracy (Pip Install ASR), while also changing the distribution of hallucination package names. Experiments show that this attack can evade existing static analysis, large language model-based, and agent-based defense mechanisms. This finding means that seemingly harmless prompts can covertly manipulate the model’s hallucination behavior, thereby creating security risks in downstream software supply chains.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
arXiv

SignalsSIGNALS

Keyword heat
  • arXiv1

All reports (1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills

研究人员提出“中性提示攻击”(NPA),一种利用语义良性指令(如鼓励想象和详尽性)增加包幻觉倾向的隐蔽攻击范式,该攻击不指定特定恶意包名而是将模型依赖生成行为转向推测性名称。研究者在多个面向代码生成的 LLM 及包幻觉基准上评估了 NPA,结果显示其提高了幻觉准确率(Hallucination ASR)和 pip 安装准确率(Pip Install ASR),改变了幻觉包名分布,并规避了现有的静态分析、基于 LLM 及基于 Agent 的技能防御。这些发现表明看似无害的提示可 covertly 操纵幻觉行为并制造下游软件供应链风险。