Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
Researchers have proposed a new type of hidden attack paradigm called “Neutral Prompt Attack” (NPA). This attack utilizes semantically benign instructions (such as encouraging imagination and detail) to redirect the generation behavior of large language models towards speculative names without specifying specific malicious package names. In evaluations against code-generation LLMs and package hallucination benchmarks, NPA significantly improved hallucination accuracy (Hallucination ASR) and pip installation accuracy (Pip Install ASR), while also changing the distribution of hallucination package names. Experiments show that this attack can evade existing static analysis, large language model-based, and agent-based defense mechanisms. This finding means that seemingly harmless prompts can covertly manipulate the model’s hallucination behavior, thereby creating security risks in downstream software supply chains.