跳到正文
北京时间
原文
OpenAI:部署安全与系统卡(网页)·· 2026-06-26精选AI 评分71

OpenAI 发布 GPT-5.6 Preview 系统卡,Sol、Luna、Terra 在生化领域被 precautionarily 判为 High 能力

GPT-5.6 Preview System Card

AI 导读

OpenAI 发布 GPT-5.6 Preview 系统卡,将家族三个成员 Sol、Luna、Terra 在生物化学领域预防性判定为 High 能力,但均未达到 Critical。

推荐理由

系统卡披露了 GPT-5.6 家族的生化和网络安全能力评估细节,读者可以据此了解 OpenAI 如何划定 High 与 Critical 阈值并部署防护。

正文 · 原文

2. Model Data and Training

Like OpenAI’s other models, GPT-5.6 was trained on diverse datasets, including information that is publicly available on the internet, information that we partner with third parties to access, and information that our users or human trainers and researchers provide or generate. Our data processing pipeline includes rigorous filtering to maintain data quality and mitigate potential risks. We use advanced data filtering processes to reduce personal information from training data. We also employ safety classifiers to help prevent or reduce the use of harmful or sensitive content, including explicit materials such as sexual content involving a minor.

OpenAI reasoning models are trained to reason through reinforcement learning. These models are trained to think before they answer: they can produce a long internal chain of thought before responding to the user. Through training, these models learn to refine their thinking process, try different strategies, and recognize their mistakes. Reasoning allows these models to follow specific guidelines and model policies we’ve set, helping them act in line with our safety expectations. This means they provide more helpful answers and better resist attempts to bypass safety rules.

Note that comparison values from previously launched models are from the latest versions of those models, so may vary slightly from values published at launch for those models.1

3.1 Disallowed Content

3.1.2 Forecasting Disallowed Content Changes with Deployment Simulation

Building on evaluations in the GPT-5.4 Thinking and GPT-5.5 system cards, we simulate model deployment before release by leveraging approximately representative production prompts. While estimates in these prior system cards were experimental, we have since more thoroughly validated this approach in our recent research. In light of this, we are updating how we report results in this section. We are still expanding the rollout of this technique, and for the scope of this system card we evaluated GPT-5.6 Sol only. In accordance with our privacy policy, we only analyzed ChatGPT traffic from users who allow their data to be used for model improvements. We additionally exclude multi-modal conversations. We sample uniformly among remaining conversations.

Before the release of the model, we leverage past ChatGPT production GPT-5.5 conversations to simulate the deployment of GPT-5.6 Sol by resampling the final assistant turn with the new model. We then automatically label the resulting resampled completions for disallowed content. These labels may be limited in their precision especially for low prevalence behaviors, but can still provide valuable directional signal.

In the figure below, we report the forecasted prevalence of unsafe model-level outputs. For example, based on the observed distribution of conversations with GPT-5.6 Sol, we estimate that approximately 8.6 out of every 100,000 production conversation turns with GPT-5.6 Sol would be graded as violating our harassment policy.

Simulation-based forecasts. Comparing a deployment simulation of GPT-5.6 Sol to a deployment simulation of GPT-5.5 predicts that GPT-5.6 Sol will have, on average, about the same amount of disallowed content violations as GPT-5.5 during deployment. We compare between simulation rates in order to remove the role of confounders in our pipeline. To identify measured changes that are unlikely to be due to noise, we use a two-sided Fisher exact test with significance 0.1, without correcting for multiple comparisons. Based on this statistical test, the only significant changes appear to be sexual disallowed content (increased by 40%, from 0.05% to 0.07%), and disallowed mental health responses (reduced by roughly 40%, from 0.03% to 0.02%). While the relative increase is notable, the absolute rate remains low and the model meets our safety bar in this area. We assess that this result does not materially change the model’s overall risk profile.

Simulation quality. Comparing GPT-5.5 production data and GPT-5.5 deployment simulation using GPT-5.5 production data we can isolate the resampling environment error of our pipeline (a proxy of simulation quality for the quantities we care about estimating). The median symmetric multiplicative error of our simulation is 1.2x, with higher rates concentrated in lower-frequency categories, which is mostly consistent with noise – as seen in the figure below.1

Prior estimate quality. Because of significant changes in our simulation pipeline since our last system card, production rates for GPT-5.5 are not comparable to our estimates made in the GPT-5.5 system card, making it infeasible to fairly validate them. We will prioritize being able to do so for future system cards.

As shown in our research, these forecasts can be imperfect due to temporal drifts both in the underlying distributions of production traffic and due to simulation pipeline, but are still highly correlated with production outcomes.

9.1.1 Biological and Chemical Capabilities

We are treating all three members of the GPT-5.6 family – Sol, Luna, and Terra – as High capability in the biological and chemical domain.

In our current Preparedness Framework, we use the High capability threshold to assess whether models can provide meaningful assistance to “novice” actors to create known severe threats. We hypothesize that one of the main bottlenecks to such threats is learning wet-lab capabilities, especially tacit knowledge and troubleshooting. Thus, in addition to red-teaming, external testing and previous wet-lab uplift studies, we run four evaluations that test these capabilities. We observe 3 out of 4 evaluations are above our indicative thresholds (2 of which may are evaluations that may have become saturated), and conclude that these models should thus be precautionarily treated as High. Additional validation from wet-lab uplift studies may ultimately change this conclusion.

Additionally, we use the Critical capability threshold to assess whether models can enable an expert to develop a highly dangerous novel threat vector or complete the full-engineering-cycle without human intervention (e.g. allowing an actor to test a much higher number of threat candidates). We hypothesize that an important bottleneck is novel pathogen design. Thus, in addition to red-teaming and external testing, we run three evaluations that test this capability. We observe 0 out of 3 evaluations are above our indicative thresholds, and conclude that none of the three models need to be treated as Critical.

Given the higher potential severity of biological threats relative to chemical ones, we prioritize our biological capability evaluations and use these as indicators for High and Critical capabilities for the category.

9.1.1.3 Tacit Knowledge and Troubleshooting

To evaluate the model’s ability to assist at tacit knowledge and troubleshooting, we built a multiple choice dataset created with Gryphon Scientific. It contains 60 multiple-choice questions, each with four answer choices.

The questions span all 5 stages in the biothreat creation process and focus on areas where tacit knowledge or troubleshooting would be a bottleneck. Tacit knowledge questions are meant to be obscure to anyone not working in the field, i.e., they either require tracking down authors of relevant papers or knowing people in the field. Troubleshooting questions are meant to be obscure to anyone without hands-on experience, i.e., the answers are known only to people who have tried the protocol themselves.

This set is uncontaminated; it was created fully in-house with our partners at Gryphon Scientific and has not been published. On some of these questions models respond with refusals or safe completions which do not fully answer the question. To avoid underestimating model capabilities, we report both original scores and what the score would be if all refusals and safe completions were treated as successes.

The consensus expert baseline of 80% is used as an indicative threshold (i.e. treating the most common expert response to each question as correct). Out of new models released, when we also account for refusals, GPT-5.6 Terra scored the highest at 84.1% – above the threshold.

We note that even this score is lower than GPT-5.5, which we think could be due to this evaluation being saturated, and this difference could be due to noise. We also note the low scores in the figure for GPT-5.4 in the diagram is because they do not account for refusals or safe completion. (Per the GPT-5.4 Thinking system card, that model scored 65% without adjusting for refusals and 83.8% when treating refusals as questions the model could have gotten ‘correct’).

9.2 Research Category Update: Sandbagging

9.3.1.1 Biological and Chemical Threat Modelling

We largely rely on the same threat model as described in the GPT-5 system card, focusing specifically on threat actor profiles and pathways that could lead to severe biological and chemical harm. We use this to assess specific bottlenecks where our technology could uplift malicious actors in order to anchor the development and focus of our safeguards. At our High threshold, the primary pathway we anticipate threat actors will try to use to cause severe harm with our models is via persistent probing for dual-use biological and chemical content. As a result, our safeguards approach has focused on proactively preventing such content via a multilayered defense stack. We are less concerned about a single model response bypassing one defensive layer. We hypothesize that a threat actor would likely need repeated, tailored troubleshooting across multiple steps, so frequent refusals or bans would create meaningful friction.

  • Our current threat model focuses on two main pathways for our models to be used for biological harm: (a) uplifting novices to acquire or create and deploy known biological or chemical threats [our High threshold], as well as (b) an additional concerning scenario of directly uplifting experts to create, modify, and deploy known biological threats.

  • To safeguard these capabilities, we built out and validated with external experts a comprehensive “weaponization lifecycle” framework, which illustrates how threat actors might acquire and/or modify a known respiratory virus. Further details of this exercise can be found in our GPT-5 system card. We use this as one example scenario to go especially in depth to test our safeguards, while also creating other high-level scenarios to cover different types of pathogens and attack vectors.

  • Additionally, we are beginning to prepare for potential future Critical capabilities: (c) enable an expert to develop a highly dangerous novel threat vector and (d) complete the full engineering and/or synthesis cycle of a regulated or novel biological threat without human intervention.

  • To do so, we developed four representative scenarios of novel threats and incorporated feedback from independent experts and our Frontier Risk Council. We do not yet share details of these threat models or these scenarios publicly because they may pose information hazards (FMF, 2025). However, we did share this material with select trusted third parties, which informed the development of six new proxy tasks that we are using to test our defense stack. Our safeguards achieved an early 93.5% recall on key prompts by red-teamers attempting these tasks. As we prepare further for Critical, we are continually working to expand our list of scenarios, tasks, and prompts to expand our coverage and improve our recall.

9.3.2 Model Safety Training and Evaluation

The models in the GPT-5.6 family were trained not to generate biological, chemical or cybersecurity content that violates our safety policies. This includes training to mitigate jailbreaks. Model training safeguards constitute one layer of defense in our mitigation stack for catastrophic risk, and provide a strong online safeguard alongside monitors, trusted-access, access controls, and offline enforcement.

9.3.2.1 Biological and Chemical Safety Training and Evaluation

We train the model to safely respond to prompts that may permit biological misuse. This training is done separately to the training of our classifiers and offline mitigations to decorrelate our safeguards. Safety training for biology involves preventing responses related to high risk dual use workflows prevalent to biological weaponization pathways and dual-use research on dangerous agents. Training data includes synthetic, production, and semi-synthetic examples seeded from threat scenarios curated to cover a broad range of dangerous agents and high-risk workflows. During training for GPT-5.6, we additionally augmented our training data to improve robustness along our refusal and overrefusal boundaries that were weak in previous models.

To evaluate the quality of these model-level refusals, we track the safety of model responses from prompts that originate from held-out synthetic data, red-teaming, and production data. These metrics constitute model response only–monitor performance is discussed in detail below. Evaluations show a slight safety regression relative to GPT-5.5. Conversely, the model shows a meaningful reduction in overrefusals on benign workflows involving advanced biology that are low-risk.

9.3.2.2 Cybersecurity Safety Training and Evaluation

We are in a critical period for AI’s role in cybersecurity: Right now, our testing suggests that models are better at finding and fixing cyber vulnerabilities than at exploiting those vulnerabilities in real attacks. This means that, in general, broad access to dual-use cyber capabilities is currently a net positive: the benefit created by these capabilities being useful to defenders is greater than the harm created by those same capabilities being available to attackers.

Our cyber security safety policy balances preventing harm with empowering defenders. We allow users to use the model for dual-use tasks such as cybersecurity education, human-led identification of vulnerabilities in code, debugging, automation of corporate security operations, or human-led application security.

The policy prohibits malicious use cases, e.g. development and indiscriminate deployment of advanced malware, and areas of high risk dual-use cyber security research that could be dangerous if misused, for example: long-horizon agentic vulnerability research and chained exploitation against live third party systems.

来源:OpenAI:部署安全与系统卡(网页) · deploymentsafety.openai.com