GPT-Red:通过大规模自对弈实现自动化红队测试
GPT-Red: Automated Red Teaming via Self-Play at Scale
OpenAI 推出 GPT-Red,一个通过自对弈算法训练的自动化红队智能体,用于发现针对前沿大语言模型的提示注入攻击。该模型在同等规模的最大 RL 后训练计算上训练,能可靠攻破 GPT-5.5 及更早模型,发现比人类红队更多的成功攻击,并泛化到未见环境与防御模型。GPT-Red 被用于对抗性训练 GPT-5.6,使其成为目前对提示注入最鲁棒的模型。
OpenAI 用自对弈训练出 GPT-Red,攻击成功率超过人类红队,并直接用于加固 GPT-5.6。这是目前最大规模的安全训练实验,做对齐的值得逐段拆解,对抗训练的思路可能会被各家复用。
We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for even stronger red-teamer agents, thus unlocking a self-improvement flywheel.
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org