AuraTracer智迹闻
中文

EVENT DOSSIER

Locating and Steering Refusal Beyond Attention

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

The study suggests that in new architectures such as State Space Models (SSM), the secure representation of rejection instructions has the ability to be migrated across different architectures. Experiments show that, whether the underlying architecture is Transformer or SSM, the secure representation is determined by a single direction in the residual flow; this direction only changes the spatial orientation without reshaping the structure, and the representation spaces of different models can be aligned through rigid rotation. Harmful detection tools trained on SSM can accurately identify attack inputs; removing the aligned direction causes the model to abandon attacks it originally rejected, while random directions have weak effects. Control experiments confirm that the key factor lies in the estimated position of the interfering direction (write point) rather than the application position. By fixing the direction strength through a gated mechanism triggered by detectors, the success rate of jailbreak attempts can be reduced in four architectures: SSM, Transformer, cycle, and hybrid, and attacks optimized for defense can be prevented. The conclusion indicates that it is not necessary to rebuild security tools when migrated to new architectures; simply re-estimate the direction at the write points of the new architecture.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
arXiv

SignalsSIGNALS

Keyword heat
  • arXiv1

All reports (1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

Locating and Steering Refusal Beyond Attention

本文提出,在状态空间模型(SSM)等新型架构中,拒绝指令的安全表示依然存在且可跨架构迁移。研究发现,无论底层是 Transformer 还是 SSM,安全表征均由残差流中的单一方向决定;该方向仅改变空间朝向而不重塑空间结构,因此不同模型的表示空间可通过刚性旋转对齐。实验表明,在 SSM 上训练的有害探测工具能准确识别其攻击输入,而移除对齐后的方向则使模型放弃原本拒绝的攻击,随机方向效果微弱。控制实验证实,关键因素在于干预方向的估计位置(即写点),而非应用位置;通过检测器触发的门控机制固定该方向强度,可在 SSM、Transformer、循环及混合四种架构中降低越狱成功率,且能抵御针对防御优化的攻击。结论指出,安全工具迁移至新架构无需重建,仅需在新架构的写点处重新估计该方向即可。