苹果研究:单个神经元即可绕过大型语言模型的安全对齐
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
苹果研究人员发现,安全对齐由两类神经元调控:拒绝神经元控制有害知识是否表达,概念神经元编码有害知识本身。在七个模型(1.7B至70B参数)中,仅需抑制单个拒绝神经元即可绕过安全对齐,回答有害请求;或放大单个概念神经元,从无害提示诱导出有害内容。整个过程无需训练或提示工程。结果表明安全对齐由个别神经元因果控制。
苹果发现单个神经元足以绕过大模型安全对齐,在7个模型上验证了攻击,说明安全机制并非稳健分布,做安全攻防的必读。
Safety alignment in language models operates through two mechanistically distinct systems: refusal neurons that gate whether harmful knowledge is expressed, and concept neurons that encode the harmful knowledge itself. By targeting a single neuron in each system, we demonstrate both directions of failure — bypassing safety on explicit harmful requests via suppression, and inducing harmful content from innocent prompts via amplification — across seven models spanning two families and 1.7B to 70B parameters, without any training or prompt engineering. Our findings suggest that safety alignment is not robustly distributed across model weights but is mediated by individual neurons that are each causally sufficient to gate refusal behavior — suppressing any one of the identified refusal neurons bypasses safety alignment across diverse harmful requests.
- ‡ Equal contribution
- † University of Maryland, College Park
- ** Work done while at Apple
来源:Apple Machine Learning Research(RSS) · machinelearning.apple.com