摘要

OpenAI 发布《Path to Astra: critical capabilities and frontier safeguards》,宣布 Astra 将是其在 Preparedness Framework 下首个达到 Critical 网络安全能力阈值的模型,并给出发布所需的更强护栏细则。框架吸收了 Hugging Face 事件(ev-20260818-04)的教训,覆盖攻击者驱动与失效驱动两类威胁建模、agentic 模型接触关键系统的风险映射,以及随时间跟踪 cyber 能力的行动阈值。

为什么重要
这是前沿实验室第一次公开把模型推到自家 Critical 层级,同时说清运营上会改变什么:监控义务、沙箱深度、前沿 cyber 部署的披露责任。对安全与平台团队,这是高能力 agent 未来合规面的早期模板,也是迄今最明确的信号——Astra 的上线门槛取决于护栏部署,而不是模型能力本身。
技术细节
要点 Astra 是 OpenAI 首个在 Preparedness Framework 下达到 Critical 网络安全能力阈值的模型
框架要素 威胁建模(攻击者驱动 + 失效驱动)、agentic 模型与关键系统接触点的风险映射、随时间跟踪 cyber 能力的行动阈值
承接 HF 事件 直接吸收 Hugging Face 事件与 RL pacing 的教训(ev-20260818-04):CoT 监控分层、分级评测规则、生产 harness 加固
状态 治理框架文章;未公布 Astra 发布日期
后续更新
2026-09-03 Framework executed: GPT-6 Astra launched 9/3 (ev-20260903-01) as the first model to meet the Critical cyber threshold — the launch version ships safeguards that refuse advanced-cyber tasks (e.g. PoC exploit creation), with Daybreak expanding access under less restrictive safeguards in the coming weeks; OpenAI also discloses Astra's written reasoning is harder to monitor than GPT-5.6 Sol's (fewer written steps), with misalignment monitoring live in production for Astra-class models. The 'launch gated on safeguard deployment' reading held
标签
openaiastrapreparedness-frameworkcritical-capabilityfrontier-safeguardsgovernance