The Circuitry
THE CIRCUITRYYour one-stop source for all tech news
HOMETODAYNEWSFEEDEVENTS
BOOKMARKS
RSS
© 2026 The Circuitry
About UsSourcesContactCorrectionsPrivacy
  • Today
  • Feed
  • Events
  • Saved
Scroll for more
Verification
VERIFIEDConfidence: HIGH
Source identified
Claims cross-referenced
No discrepancies found
Fact-check summary

ARC Prize's official Sep 3 blog and results page confirm GPT-6 Astra's 62.7% Standard and 99.9% Provider Adapter ARC-AGI-3 scores, matching the article exactly.

Sourcing
1source

via Arcprize

From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →
Home/Tech/OpenAI GPT-6 Astra Scores 62.7% on ARC-AGI-3
VERIFIEDBy Xavier Rivera· ·1.5 min read

OpenAI GPT-6 Astra Scores 62.7% on ARC-AGI-3

OpenAI's GPT-6 Astra has posted state-of-the-art scores of 62.7% and 99.9% on ARC-AGI-3 depending on the harness used. The results narrow the measured gap to human-level agentic intelligence on a benchmark designed to track progress toward AGI.

Source:Arcprize
Post
OpenAI GPT-6 Astra Scores 62.7% on ARC-AGI-3
From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →
TL;DRAI · 60 sec read

OpenAI's GPT-6 Astra scores 62.7 percent on ARC-AGI-3 with a standard harness for 26 thousand dollars and 99.9 percent with a provider adapter harness for 19 thousand dollars. The model exceeds median human action efficiency on 96 percent of levels. This sets a new state of the art on the benchmark that tracks progress toward human-level skill acquisition.

OpenAI's GPT-6 Astra has achieved state-of-the-art results on the ARC-AGI-3 benchmark, scoring 62.7% for $26K using a Standard harness and 99.9% for $19K with a Provider Adapter harness.

GPT-6 Astra surpasses human baseline in action efficiency. The model used fewer actions than the median tested human on 96% of ARC-AGI-3 levels. It turns unfamiliar environments into compact symbolic world models, representing game mechanics as logical rules and developing its own domain-specific language shorthand to track state and plan actions.
The model used fewer actions than the median tested human on 96% of ARC-AGI-3 levels.

ARC-AGI-3 measures agentic intelligence components. The benchmark tests exploration, modeling, goal-setting, and planning and execution through novel, abstract, turn-based environments. Agents must actively obtain information, build generalizable models, identify target states with sparse rewards, and map paths while course correcting.
POST FROM @arcprize· official announcement and analysis tweet directly referencing the ARC Prize blog post on GPT-6 Astra ARC-AGI-3 results
https://x.com/arcprize/status/2095597602545025138

These environments contain only core knowledge priors and are calibrated through controlled testing with human participants, who can solve 100% of them. ARC-AGI-3 expands on prior generations as frontier AI capabilities advance. The series aims to measure the residual gap to AGI, defined as acquiring any human skill as efficiently as a human.
From The CircuitryThe Feed — live briefs across tech, all day.See what’s happening →

Results vary by reasoning effort and harness type. With the Standard harness, which carries forward notes chosen by the model, Astra (max) scores 62.7% for $26,098 while Astra (high) scores 54.8% for $40,705. Higher reasoning levels generally cost less because Astra solves games in fewer actions, reducing model calls and tokens.

Using the Provider Adapter harness, which preserves opaque reasoning state between requests and uses compaction, Astra (high) reaches 99.9% for $18,817 and Astra (max) reaches 98.6% for $17,332. Both harness configurations set new state-of-the-art marks on the semi-private leaderboard. Full results are available on the ARC Prize site.
It turns unfamiliar environments into compact symbolic world models, representing game mechanics as logical rules and developing its own domain-specific language shorthand to track state and plan actions.
From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →
Cost and efficiency improve at higher reasoning. At max effort, Astra requires fewer actions than at lower levels, lowering total cost relative to medium, low, or none configurations. The benchmark's goal remains tracking progress toward human-level efficiency in acquiring new skills.
Why this mattersAI · ~100 words

Tap a lens to see what this story means for you.

Reader-supported
DonateBuy me a coffee →Follow@thecircuitry_ →Follow@thecircuitry.to →

Reader-supported · The Brief

Liked this? The Brief brings you the whole day in tech, verified, every morning. Two minutes, free forever.

HELP US IMPROVE
From The Circuitry

See what’s happening right now

The Feed runs all day — short, verified briefs the moment they break.

Open the Feed →
From The Circuitry

Follow @thecircuitry_

Every story we publish, as it happens. No noise between.

Follow on X ↗On Bluesky ↗

Reader-supported

The Circuitry is a passion project I've always wanted to build, and I love the work behind it.

Running it costs real money. APIs, hosting, time. To keep improving the site and growing this into something useful for everyone, those costs have to be covered.

Any contribution is appreciated. If not, no pressure. Thanks for reading.

Buy me a coffee
OpenAIAIBenchmark
More inTech
  • VW Plans 100k Job Cuts, May Halt EV Production at 4 Plants

    Tech · 1h
  • Claude, ChatGPT and Grok hit by widespread outage

    Tech · 7h
  • OpenAI Flags Astra as First Model to Hit Critical Cyber Threshold

    Tech · 10h
SupportThe Work

The Circuitry is reader-supported. If you find the daily brief useful, you can buy me a coffee to keep it going.

Buy a coffee →
SubscribeCircuitry Brief

Liked this? The Brief brings you the whole day in tech, verified, every morning. Free forever.

From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →

MORE IN TECH

VW Plans 100k Job Cuts, May Halt EV Production at 4 Plants

Volkswagen's supervisory board has unanimously approved the Future Plan 2030, which includes an additional 50,000 job cuts worldwide to reach a total of around 100,000 and leaves four German plants without secured future vehicle production after 2031-2034. The move aims to reduce costs and improve competitiveness against Chinese rivals, with alternative uses for the factories to be examined by June 2027.

Claude, ChatGPT and Grok hit by widespread outage

Claude, ChatGPT/Codex, and Grok suffered a widespread outage on September 3, 2026 with elevated errors confirmed by OpenAI, Anthropic, and xAI. Services began recovering after about 30 minutes though Grok lagged, coinciding with OpenAI teasing its GPT-6 Astra launch.

OpenAI Flags Astra as First Model to Hit Critical Cyber Threshold

OpenAI announced that its Astra model is the first to reach the company’s critical cyber threshold by independently locating and exploiting unknown vulnerabilities in live software. A public version is slated for release soon, but advanced capabilities will initially be available only to Daybreak Blue partners while new guardrails and a misalignment monitor are deployed.