OpenAI 用 AI 攻击自家 AI:GPT-Red 自动发现安全漏洞,成功率 84% 远超人类
OpenAI is now using AI to attack its own AI, and it's working better than humans ever did
OpenAI 训练了内部 AI 模型 GPT-Red,通过自我对弈强化学习自动模拟提示词注入等攻击,在测试场景中成功率达 84%,而人类红队仅为 13%。GPT-Red 的发现直接用于训练,使 GPT-5.6 Sol 在直接提示词注入上的故障次数比四个月前的最佳模型减少六倍,且未影响通用性能。约 3.8% 的“更强”提示词注入仍能成功,GPT-Red 暂不对外开放。
OpenAI让AI攻击自己找漏洞,成功率84%把人甩在后面,直接拉低了GPT-5.6的注入失败率六倍,这比人类红队靠谱多了,但别忘了3.8%的缺口照样能捅娄子。
OpenAI trained an internal AI model called GPT-Red to automatically find security flaws in GPT models. GPT-Red simulates prompt injections and other attacks where malicious instructions hide in emails, websites, or files. Trained via self-play reinforcement learning, GPT-Red attacks while defender models block, and both improve over time. It finds successful attacks in 84 percent of test scenarios versus 13 percent for human red teamers. In one test, it manipulated an AI-powered vending machine in OpenAI's office, changed prices, and canceled other customers' orders.
The results feed directly into training. GPT-5.6 Sol shows six times fewer failures on direct prompt injections than the best model from four months ago, OpenAI says, without hurting general performance. But about 3.8 percent of "stronger" prompt injections still succeed. Scale that to hundreds or thousands of attempts, and a sizable number get through, similar to Claude Opus 4.5.

GPT-Red stays internal; a paper with more details will follow.
来源:The Decoder:AI News(RSS) · the-decoder.com