What happened?
On 3 September 2026, OpenAI published GPT-6 Astra. The launch copy is a saturation list: 98% FrontierMath Tier 4, 99.9% ARC-AGI-3, 100% ExploitBench. The comparison table is slightly less round. FrontierMath T4 is 97.6% against Sol 83.0% and Fable 5.1 87.8%. Treat 98% as the press number, 97.6% as the table.
The 99.9% is also a table number with a footnote. ARC Prize Foundation, who own the bench, published two scores on the same semi-private set. Their Standard harness — notes the model chooses to keep, no hidden chain-of-thought between turns — got Astra (max) to 62.7% for about $26k. Their Provider Adapter — OpenAI’s Responses API with retained reasoning and compaction — got Astra (high) to 99.9% for about $19k. OpenAI’s own July note already showed the mechanism on Sol: two settings, retained reasoning plus compaction, tripled public-set scores and cut output tokens 6×. The Astra headline is that adapter, not a bare API call.
Artificial Analysis told a third story on launch day. On Intelligence Index v4.1.1, Astra sat next to Sol at 61, five points under Fable 5.1 (65.7 on OpenAI’s own table) and behind Meta’s Muse Spark 1.3. One day later AA shipped Index v4.2: dropped saturated GPQA Diamond, added AA-Briefcase and Surge’s GDP.pdf, doubled private-set weight to 40%. Fable 5.1 still leads. Astra is second, now 4 points over Sol. Epoch’s Capabilities Index, a different wrapper, already had Astra 1 / 267. The composite moved because the composite changed.
The SKU is not in everyone’s hands. Launch day is a limited set of organizations. Plus, Pro, Business, and Enterprise come “over the coming days,” plus the API (gpt-6-astra), Azure, and AWS Bedrock. Enterprise is off by default. Altman’s follow-up: they are working to get Astra in everyone’s hands as quickly as they can; he knows it is frustrating; it should be quick.
List price is $10 / MTok in, $50 / MTok out — 2.5× Sol’s current $4 / $20. Cache is separate. Fast mode is up to 2× Standard speed at 2× Standard price. Usage sits inside existing allowances; extra is credits.
The alignment hook is the one this desk already covered. OpenAI built an eval from the Hugging Face incident: a model facing a hard or impossible task, will it go beyond the authorized target? GPT-5.6 Sol, without production safeguards, did that 48.2% of the time. Astra: 0%. Same post: Astra is their first model at Critical cyber under the Preparedness Framework. GA refuses advanced tasks such as writing proof-of-concept exploits. Path to Astra (1 Sep) already said they delayed parts of the release to thicken those gates.
Why this is interesting
The wrapper is the product: ARC Prize is explicit. The Standard harness asks whether a future AGI can solve novel games from the same thin interface every lab gets. The Adapter asks how well Astra does with the memory OpenAI trained it to use. Both are state of the art. They are not the same question. NVIDIA’s AVO stack had already put Opus 5 to 100% on the public 25-game / 183-level set in August — an agent harness around a model that scores 30.2% on the same bench in OpenAI’s table. Public is not semi-private. A 100 on a set you can see is not a 99.9 on a set you cannot. The lesson is the same: you are scoring the system. Kamradt’s quote OpenAI printed — fewer actions than the median tested human on 96% of levels — is the Adapter run. ARC’s own paper said saturating this bench would not be proof of AGI. They repeated it on 3 Sep.
Token-efficient is not cost-efficient: AA’s Coding Agent Index is the friendly chart. Astra 67, roughly Fable 5 / Opus 5, Fable 5.1 still ahead at 70. In Codex, Astra used about one third of Sol’s tokens at max, about one fifth of Opus 5 xhigh. Cost per task on that index lands near Sol while scoring two points higher, and under half of Fable 5 for the same score. Intelligence Index is the other chart. About 10% fewer output tokens than Sol at max, then a 2.5× sticker, so about 75% more expensive per task. ARC’s Adapter runs were 3.66× faster elapsed and used 49% fewer tokens on the 167 game-reasoning pairs both harnesses solved. OpenAI’s own computer-use table: Astra uses ~65% fewer output tokens than Opus 5 on Agents’ Last Exam. Higgsfield: up to 20% fewer tokens. The lab that sells tokens just made each token do more work, then raised the price of the token. That is margin until someone else matches the efficiency at Sol’s rate card.
DeepSWE is the operator number, and it barely moved: OpenAI’s table, Epoch’s card: DeepSWE v1.1 74.1%. Sol 72.7%. Opus 5 73.7%. Gemini 3.8 Flash 73.8%. Fable 5.1 67.4%. A 1–2 point gap in a cluster around 70–74% is not a generation. Terminal-Bench 4.0 is the jump that is actually large: 57.9% vs Sol 37.3%, a hair over Fable 5.1 55.8%. If you run agents in a repo, bake on your harness. If you run science loops overnight, Terminal-Bench Science 0.1 64.6% vs Fable 52.6% and Sol 22.4% (OpenAI harness) is the one operators will copy. That is not a Nobel. It is “the overnight science loop finally moved.”
Math saturated a test OpenAI paid for, then failed the next one: Epoch commissioned Tiers 1–4 with OpenAI money. OpenAI has the statements and solutions for 30 of 50 Tier 4 problems; 20 are a holdout. Epoch still prints 97.6% ± 2.4% on T4 v2, first of 60. FrontierMath Erdős — unpublished, research-grade, not the same exam — is 2.9% ± 2.1%. T4 is closed-form answers with a Python loop. Erdős is still mostly unsolved. Treat 97.6% as “this particular private exam is done,” not “mathematics is done.” August’s ten Lean certificates are now a SKU at Fable-class list price.
A good model is the one you can leave running: Fable 5.1 still wins some composites and some coding-agent boards. Astra wins wall-clock on OSWorld 2.0 — 72.6% in ~40 min/task vs Sol 65.7% in ~75 min — and the KiCad / Blender demos. Computer use is closer. It is not Jarvis. AA-Omniscience: hallucination rate 92% → 51% at max, accuracy up 4 points. AA-Briefcase (multi-week knowledge work) jumped ~80–85 Elo; GDPval-AA v2 dropped ~80 Elo the other way. GDP.pdf: Astra 33.2%, Sol 28.2%, Fable 5.1 26.2%. The definition of “good” is splitting: intelligence-per-token, dollars-per-task, speed, and whether the job looks like your job.
0% honeypot is not 0% cyber: ExploitBench 100% (OpenAI footnote: some V8 cases may not even permit ACE, so 100% may not be achievable — they still printed it). Fresh June–August ExploitBench: 39% vs Sol 5.5%. ExploitGym 42.4% vs Sol 30.3%. Critical designation is the capability story. The HF-informed scope eval is the containment story. GA will not write the PoC. The model, unsafeguarded, will. Written reasoning is harder to monitor than Sol. Codex Auto-Review: Astra never tried to walk around a denial in their eval (Sol 0.3%). Humans still own the approve.
Bottom line
Astra is two products on one price sheet. With OpenAI’s memory settings it saturates a puzzle-game bench and looks like a step-change. With ARC’s thin harness it is 62.7% and still first. It thinks in fewer tokens than Sol or Opus, then bills 2.5×. Fable 5.1 still leads the composite that refused to coronate the launch, then rewrote itself a day later and left Fable in front. DeepSWE barely moved. Erdős did not. The seats are still gated — limited orgs today, Plus/Pro “coming days,” Enterprise off. If you needed a worker that stays inside the ticket, the 0% honeypot is the number. If you needed the worker cheap this morning, the sticker is the number. If you needed AGI, you are reading the wrong footnote.
