What GPT-6 Astra is and what it can do

Astra was described by OpenAI as "the most intelligent and aligned model we have ever built." The focus is not just reasoning — it is autonomous execution. The company positions the model to operate computers the way a person would: navigating interfaces, executing code, conducting scientific research, and producing professional documents without step-by-step human intervention. The launch campaign tagline summarizes the positioning: "Anything you can do on a computer, Astra can do for you. Fast." The model was trained in the largest compute run in OpenAI's history — more than 100,000 GPUs at the Stargate complex in Texas — and is the first where earlier OpenAI models actively supervised the training of the new one.

Benchmarks: real advances and declared controversies

The numbers published by OpenAI are aggressive. Astra reached 97.6% on FrontierMath Tier 4, an advanced mathematical benchmark; 99.9% on ARC-AGI-3 using OpenAI's internal harness; and 100% on ExploitBench, which measures the ability to turn known vulnerabilities into functional exploits. On computer use, the model scored 72.6% on OSWorld 2.0, completing tasks in approximately 47% less time than its predecessor GPT-5.6 Sol. But controversies arrived alongside the results. The ARC Prize — the independent organization that administers ARC-AGI-3 — reported 62.7% using its standard harness, far below OpenAI's 99.9%. The difference lies in the method: OpenAI used a "Provider Adapter" that preserves reasoning state between calls, a technique unavailable in the standard harness. The result is impressive either way — but direct comparison with other models is compromised.

Cybersecurity: OpenAI's first "Critical" model

Astra is the first OpenAI model to reach the "Critical" level on its own Preparedness Framework — the company's internal risk assessment scale for cybersecurity. In practice, this means the model has the capacity to execute tasks that could cause significant harm to digital infrastructure. On ExploitBench with recent vulnerabilities — June through August 2026 — Astra scored 39% against GPT-5.6 Sol's 5.5%. On SRE-Bench, which measures reverse engineering of binaries without access to source code, the model solved 88% of tasks on the first attempt and 99.2% within four attempts. For this reason, the most sensitive cybersecurity capabilities are not publicly available — they are restricted to vetted organizations in the Daybreak Access program. The public model refuses tasks such as generating proof-of-concept exploits.

Controversial rollout and pricing 2.5 times higher

The launch was not without turbulence. Access began restricted to a limited group of organizations in the Daybreak program, leaving paying subscribers — including Pro users — without access on the first day. Sam Altman had to publicly apologize and guaranteed priority access for Pro subscribers in the next rollout phase. API pricing is $10 per million input tokens and $50 per million output tokens — 2.5 times the price of GPT-5.6 Sol. The context window is 1,050,000 tokens, with a maximum output of 128,000 tokens. The model is also available via Amazon Bedrock and Microsoft Foundry. The knowledge cutoff is April 30, 2026.

A week of race to the top

The launch of GPT-6 Astra comes two days after Anthropic launched Claude Fable 5.1 on September 1, 2026. The week also concentrated announcements from Meta and Google. Sam Altman described the moment as "a new generation of entrepreneurship." Brockman went further: he stated that he personally believes OpenAI has achieved AGI with this model, leaving the final interpretation to each user. Astra competes directly with Anthropic's Claude Fable 5.1 on computer use and coding benchmarks — a segment where the two models now share the top of the table.