iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Not on the published evidence assessed here. OpenAI reports that GPT-6 Astra performs strongly across several important benchmarks, but those results do not establish broad, reliable general intelligence. The most striking figure—99.9% on ARC-AGI-3—also varies substantially with the evaluation harness. That is an evidence-based conclusion, not proof that Astra lacks intelligence or that future evidence could not change the assessment.
What did OpenAI mean by “the AGI era”?
OpenAI’s launch announcement presents Astra as a major advance in computer use, browsing, software engineering, cybersecurity, science and professional work. The phrase “Welcome to the AGI era” is attributed to OpenAI President Greg Brockman in the AGI Society’s review. But the review says the statement did not define AGI or specify a test or threshold the model had met.
A claim that an era has begun is not the same as a published demonstration that a particular model satisfies an agreed standard for artificial general intelligence. The distinction matters because there is no generally accepted empirical test for AGI, according to the AGI Society review.
What do Astra’s reported benchmark results show?
OpenAI reports high scores across tests of different kinds. These are results on named benchmarks—not measurements on a universal intelligence scale. The scores below are OpenAI’s reported results from its 2026 announcement unless another source is identified.
#1 Best Overall
| Benchmark or comparison | Reported result | What the result establishes—and what it does not |
|---|---|---|
| ARC-AGI-3 | 99.9%; OpenAI says it used its Responses API harness. | A striking result on this interactive benchmark under that evaluation setup. The score is not independent of the harness used. |
| FrontierMath Tier 4 | 98%. | A strong result on this named mathematics test; it is not by itself evidence of general ability across domains. |
| OSWorld 2.0 | 72.6% for Astra versus 65.7% for GPT-5.6 Sol. OpenAI reports a latency simulation taking roughly 40 minutes per task for Astra and 75 minutes for GPT-5.6 Sol. | OpenAI reports the comparison on an offline task subset. The score and timing apply to that setup, not every computer-use task. |
| AutomationBench | 41.4%. | One result on one named benchmark; the announcement’s benchmark procedure and settings determine how it should be interpreted. |
| Terminal-Bench 4.0 | 57.9%. | A score on a particular terminal-use benchmark, not a measure of unrestricted software or computer competence. |
| Agents’ Last Exam | 59.3% for Astra, versus 53.6% for GPT-5.6 Sol and 55.5% for Claude Opus 5 in OpenAI’s comparison table. | Astra leads the listed comparison on this test; the result remains specific to the benchmark and configurations shown. |
| Terminal-Bench Science 0.1 | 64.6% for Astra versus 52.6% for Claude Fable 5.1 at the configurations described by OpenAI. | A comparison on this science-focused test, not evidence of general scientific competence in every setting. |
OpenAI also lists ARC-AGI-1 and ARC-AGI-2 results. They should not be conflated with ARC-AGI-3: the tests differ, and the ARC-AGI-3 result concerns an interactive benchmark. OpenAI’s announcement includes benchmark procedures and qualifications in its footnotes, so headline numbers should be read alongside those settings.
Why does the ARC-AGI-3 harness change the interpretation?
The AGI Society review reports ARC Prize results of 62.7% for Astra with a provider-neutral harness and 99.9% with an adapter that preserves OpenAI’s reasoning state between requests. The difference makes the evaluation setup central to interpreting the headline score. The 99.9% figure should not be presented as though it were a setup-independent result.
Rank #2
There is another meaningful, but narrower, ARC-AGI-3 comparison. OpenAI quotes Greg Kamradt of the ARC Prize Foundation saying Astra surpassed the human action-efficiency baseline on 96% of levels, which he described as effectively reaching human parity on the benchmark. This is a claim about performance and action efficiency on those levels. It is not a finding that Astra meets a general, agreed AGI standard. The AGI Society review also relays ARC Prize’s caution that benchmark saturation alone does not prove AGI.
What does independent reporting add?
Live Science, reporting results from Artificial Analysis, says Astra scored 61 on the Intelligence Index—the same score as GPT-5.6 Sol. The report also says Astra fell in relative ranking on GDPval-AA v2, a workplace-task evaluation spanning 44 occupations. Those results complicate any claim that one broad benchmark picture shows a clear across-the-board leap.
The same report describes a mixed profile: improvements in token efficiency on some software-engineering tests alongside reported regressions in customer service, scientific Python programming and long-context reasoning. Taken together, the findings suggest strengths and weaknesses can coexist; neither an aggregate index nor a few strong task results settle the question of general intelligence.
What would a convincing demonstration of general intelligence need to show?
There is no consensus test to apply, so the answer depends partly on what a proposed definition of AGI requires. A serious assessment should look beyond a collection of scores and ask whether performance transfers across unfamiliar tasks and conditions, remains reliable over longer work, and holds up under evaluations that are not specific to one provider’s setup.
- Breadth and transfer: Can the system handle materially different tasks, including novel ones, rather than excel only within benchmark formats?
- Evaluation independence: Do results persist across provider-neutral and vendor-specific harnesses, with tools, prompting, scoring and model configuration made clear?
- Reliability and autonomy: Can it complete extended tasks dependably, not merely produce a strong score on a bounded test?
- Real-world performance: Do results carry over to complete occupations or physical-world tasks? The AGI Society review says comparable published Astra results were absent for autonomous driving, independently completing a household physical task, broad robotic autonomy and performing a complete occupation. Those gaps mean the evidence is incomplete; they do not mean Astra was tested and failed those tasks.
- Independent replication and scope: Are results reproducible, and can the system stay within authorized limits while acting? OpenAI’s announcement describes deployment safeguards, but benchmark performance alone does not answer every question about safety or real-world behavior.
Scores are not directly comparable unless the task set, tools, prompts, harness, scoring method and model configuration are sufficiently aligned. That is why an impressive result should be interpreted with its specific conditions, rather than treated as a portable measure of intelligence.
Does the Navier–Stokes proof show Astra made a scientific discovery?
Not according to the account in the AGI Society review. The review says the proposed Navier–Stokes proof was generated by an internal model more capable than Astra; Astra was subsequently used to formalize and verify it. The proposed proof remains subject to independent scrutiny. It should not be credited to Astra as the original discovery.
Best Value
So, has Astra demonstrated general intelligence?
The evidence supports a narrower conclusion: Astra has demonstrated substantial capabilities on specific tests, and OpenAI reports strong performance across several domains. But the reviewed public results do not establish broad, reliable general intelligence across domains and conditions. The ARC-AGI-3 harness difference and mixed independent findings are important qualifications, while the lack of a generally accepted AGI test means no single benchmark can settle the matter.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

