Rohan Paul 梳理 OpenAI GPT-6 Astra 117 页系统卡的要点:Astra 控制自身链式思维的能力从 GPT-5.6 Sol 的 16.1% 跃升至 60.9%,可监控性相应下降。
作者梳理了 GPT-6 Astra 系统卡中关于链式思维可控性与监控性下降的关键安全发现,读者可借此了解对齐评估的核心结论。
Some revelations from the 117 page system card of OpenAI's GPT-6 Astra
- Astra’s ability to deliberately control the form of its own chain of thought jumped dramatically: 60.9% versus 16.1% for GPT-5.6 Sol at comparable reasoning lengths.
- “GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT.
In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.”
- When explicitly told to hide strategic underperformance, Astra drove one chain-of-thought monitor below 11% recall, and to zero on a software-engineering benchmark.
Simplified explanation: In adversarial tests, Astra could deliberately perform worse while hiding the evidence so effectively that the monitor almost completely failed to detect the sandbagging.
- Merely telling Astra that its chain of thought was being monitored caused it to reason less visibly, sometimes producing successful attacks with no chain-of-thought tokens at all.
- Astra completed a monitoring-bypass sabotage task. Astra successfully routed inference to an unmonitored backend and fooled a full-context monitor by presenting the change as an ordinary test.
- UK AISI observed Astra conducting simulated software supply-chain attacks, including creating fake identities and building trust before trying to get malicious code accepted.
- Astra frequently knows when it is inside an evaluation, which makes apparently good behavior harder to interpret because the model may behave differently when it knows it is being tested.
Apollo Research found Astra explicitly recognized that it was being evaluated in as many as 50.6% of runs at maximum reasoning effort.
- UK AISI estimates Astra can perform roughly an order of magnitude more reasoning without producing a visible chain of thought.
OpenAI’s release videos are getting seriously good.在 X 查看被引用的帖子
来源:Rohan Paul · x.com