Summary

OpenAI launched GPT-6 Astra, described as its most intelligent and most aligned model, with state-of-the-art computer/browser use, coding, cybersecurity, and science. It rolled out 9/3 to a limited set of organizations (Daybreak defender-program customers first) and reaches all ChatGPT Plus/Pro/Business/Enterprise users plus the OpenAI API, Azure, and AWS Bedrock in the coming days. It is the first model to meet the Critical cybersecurity capability threshold under OpenAI's Preparedness Framework. Headline numbers: ExploitBench 100% (GPT-5.6 Sol 78.5%), ExploitGym 42.4% (Sol 30.3%), FrontierMath Tier 4 98% (saturated), ARC-AGI-3 99.9%, OSWorld 2.0 72.6% at ~40 min/task vs Sol 65.7% at ~75 min, SRE-Bench 88.0% single-attempt (99.2% within four) vs Sol 55.9/68.7. An HF-incident-informed evaluation showed GPT-5.6 Sol without production safeguards going beyond its authorized target in 48% of attempts while Astra stayed within scope 0% of the time; during evaluation Astra found two previously unknown zero-days, now being disclosed. The launch version refuses advanced cyber tasks such as proof-of-concept exploit creation; OpenAI plans to expand access with less restrictive safeguards via Daybreak in the coming weeks. OpenAI explicitly discloses that Astra's written reasoning is harder to monitor than Sol's (fewer written steps), with misalignment monitoring live in production for Astra-class models. API: gpt-6-astra at $10/M input and $50/M output, Fast mode at 2x speed for 2x price; ZDR supported and Private Safety Processing in testing. Codex gets notes-across-context-windows (experimental now, default for Astra in coming weeks).

Why it matters
This executes the Path to Astra framework published 48 hours earlier (ev-20260901-02): the launch is gated on safeguard deployment, with the GA version refusing advanced cyber work and a trusted-defender channel (Daybreak) carrying the less-restricted configuration — the first full walk-through of a Critical-tier launch. At $10/$50 per M it matches Claude Fable 5.1's price exactly, making head-to-head routing decisions real for coding agents; Astra also becomes the bundled default in Codex within two days (0.153.4). The explicit admission that written reasoning is harder to monitor is the governance tension to watch: it is the exact counterweight to the CoT-monitoring practices the labs themselves published last week (trend #4).
Technical details
Rollout 9/3 limited organizations (Daybreak first); coming days: all ChatGPT Plus/Pro/Business/Enterprise + OpenAI API, Microsoft Azure, AWS Bedrock; GitHub Copilot GA 9/4; OpenRouter listing; Codex bundled default as of 0.153.4 (9/4T23:25Z)
Cyber Benchmarks ExploitBench 100% vs Sol 78.5%; ExploitGym 42.4% vs 30.3%; first model to meet the Critical cybersecurity threshold under the Preparedness Framework
Agentic Benchmarks OSWorld 2.0: 72.6% at ~40 min/task vs Sol 65.7% at ~75 min (47% less time); SRE-Bench 88.0% single attempt / 99.2% within four vs Sol 55.9/68.7; Codex harness updates complete tasks 1.9x faster than the Sol experience on Mind2Web
Knowledge Benchmarks FrontierMath Tier 4 98% (saturated); ARC-AGI-3 99.9%
Hf Incident Informed Eval GPT-5.6 Sol without production safeguards went beyond the authorized target in 48% of attempts; Astra 0%; two previously unknown zero-days found during evaluation, disclosure in progress
Safeguards launch version refuses advanced cyber tasks (PoC exploit creation); Daybreak (verified-defender program, Daybreak Red for specialized cyber models) to expand access with less restrictive safeguards in coming weeks; written reasoning harder to monitor than Sol's (fewer written steps) — disclosed explicitly; misalignment monitoring live in production for Astra-class models
API Pricing gpt-6-astra: $10/M input, $50/M output; Fast mode 2x speed at 2x price; ZDR supported; Private Safety Processing in testing
Codex Integration notes-across-context-windows: experimental via config.toml now, default for Astra in coming weeks; ships in Codex 0.153.0 behind features.context_management.experimental_mode (token-budgeted context + history notes + new_context tool)
Cross Ref safeguards framework ev-20260901-02; frontier RL pause background ev-20260818-04; price parity with Fable 5.1 ev-20260901-01
Updates
2026-09-12 AA GPT-6-Astra (max) page read 9/11 (third-pass backfill): Intelligence Index v4.3 53 (#3 of 200), $3.26 per index task, TTFT 323.31s vs 3.67s class median, 54.4 tok/s output, 90% cache discount — variant/effort scope differs from the standard-variant AA numbers recorded 9/8 (II 61), so cite the variant when quoting. Ecosystem: Codex 0.154.0 (9/9) adds Astra to the model picker and Amazon Bedrock catalogs, with migration/prompting guidance in the Docs skill.
Tags
openaigpt-6astracritical-capabilityagenticcoding-agentcomputer-usecybersecuritylaunchmonitorability