Skip to content
InYourGeek
visiteur@inyourgeek — shell
compléter historique ouvrirhelp
FR
AI· 3 min read

GPT-6 Astra hits 99.9% on ARC-AGI-3 — on a harness it built itself

OpenAI shipped GPT-6 Astra on 3 September 2026 with a near-perfect ARC-AGI 3 score. Measured on the benchmark's own standard harness, the same model drops to 62.7% — and costs more.

A scoreboard reading 99.9% with a small asterisk beside it, wires and tubing running out of the back of the display into a separate box labelled with the benchmark's name.

The headline number, and its asterisk

OpenAI put GPT-6 Astra online on 3 September 2026, first for a limited set of organisations, with a rollout to ChatGPT Plus, Pro, Business and Enterprise accounts, plus the API and AWS, promised for the days that follow (the announcement). API pricing: $10 per million input tokens, $50 per million output — exactly what Claude Fable 5 and 5.1 charge. The entire competitive pitch, then, fits in a single row of a pricing table. The model will answer to gpt-6-astra once the rollout completes.

The number doing the rounds since is 99.9% on ARC-AGI 3, a benchmark released in March 2026. Except that, as Simon Willison notes, the ARC-AGI blog publishes the protocol next to the score: those 99.9% cost $19,000 and were run on OpenAI’s in-house harness, the Provider Adapter, which holds an opaque reasoning state between requests and compacts long conversations so the model can reuse its earlier work. On ARC-AGI’s default harness, the same model manages 62.7%, for $26,000. Worse and dearer: what changed between those two numbers is not the model, it’s the plumbing wrapped around it. And Claude Fable 5 has no published result on this benchmark at all, which leaves the implied comparison with nothing to stand on.

The scores that need no special arrangements

On security, the reported results are clean, and they read best against GPT-5.6 Sol: 100% on ExploitBench versus 78.5%, 42.4% on ExploitGym versus 30.3%, and 99.2% across four attempts on the binary reverse-engineering half of SRE-Bench, versus 68.7%. A genuine leap — as long as you also notice that ExploitGym is a test the model still flunks close to six times out of ten.

Long context moves too: on OpenAI’s own eight-needle benchmark, Astra posts 100% between 256,000 and 512,000 tokens, and 96.3% between 512,000 and a million. That is the sort of figure that, if it holds up outside the lab, quietly retires a problem everyone has been papering over with RAG for two years.

The leaderboard Astra doesn’t win

Artificial Analysis puts Astra at 61 on its Intelligence Index — level with GPT-5.6 Sol, five points behind Claude Fable 5.1 at maximum effort with fallback, and behind Meta’s Muse Spark 1.3. A new generation that fails to beat the old one on the index is an odd thing to attach a fresh major version number to.

The same shop’s Coding Agent Index tells a friendlier story: at maximum effort Astra sits on the cost-efficiency frontier, costing roughly what Sol costs for two extra points, and under half Claude Fable 5’s price per task at an equal score. That, not ARC-AGI, is where the sales argument actually lives.

What the system card doesn’t tell you yet

The system card is out, but most of the numbers in circulation are OpenAI’s own self-reported benchmarks, on which Astra beats Fable in the majority of cases. The post itself, titled “A new generation of intelligence,” was hard to actually read on launch day: OpenAI’s blog was throwing 500s, and readers fell back on a mirror. The Hacker News thread had no such trouble — 638 points and 369 comments within hours. Willison, who hasn’t been given access to the model yet, declines to go further.

Scoring 99.9% on equipment you brought yourself is still a score. What the next few days will settle is how much of it is left once somebody else is holding the stopwatch.

Sources (2)

Written with AI assistance from the sources cited above, then reviewed and approved before publication by Sébastien Soulier.