OpenAI Releases GPT-6 Astra, Its First Model Rated 'Critical' Under Its Own Cyber Risk Policy
OpenAI began a limited release of GPT-6 Astra on September 3, 2026 and opened it to paying ChatGPT tiers the next day, saying the model crossed the company's top internal cybersecurity risk threshold.
Two Companies, Two Days Apart, Both Racing an IPO Clock
On September 1, 2026, Anthropic released two new models, Claude Fable 5.1 and Mythos 5.1[17]. Two days later, OpenAI answered with GPT-6 Astra, released September 3 first to a small group of trusted partners, then opened to paying ChatGPT users the next day[1][2]. Both launches landed in the same narrow window, and that timing is not a coincidence — it is the story.
Astra arrived with a label no OpenAI model has carried before. The company says it is the first of its models to cross OpenAI's own "Critical" cybersecurity risk threshold, the top tier in a scale OpenAI wrote and grades itself against[8][13]. In practice, that means OpenAI's own testers found Astra could locate software bugs nobody had disclosed yet, and figure out how to exploit them, without a person walking it through each step[8]. Because of that finding, OpenAI restricted the model's most offensive capabilities and shipped a public version that refuses some cybersecurity prompts[2].
Access comes at a price: $10 per million input tokens and $50 per million output tokens, through the API, Amazon Web Services, and ChatGPT's Plus, Pro, Business and Enterprise tiers[2][11]. A token is roughly three-quarters of an English word, so a million tokens works out to about 750,000 words — close to a long novel. Feed Astra a novel-length document and it costs about $10; get a novel-length answer back and it costs about $50. The model also carries a 1-million-token context window and a knowledge cutoff of April 30, 2026[11].
The Same Table, Two Different Winners
OpenAI President Greg Brockman said at the launch that "it's not unreasonable to feel that we are now in the AGI era"[9]. Nvidia CEO Jensen Huang, whose chips trained the model, went further three days later, posting simply "AGI has arrived"[9]. AGI — artificial general intelligence — usually means a system that can match or beat humans across essentially any cognitive task, but there is no agreed technical definition or test for it. That is exactly what critic Gary Marcus pointed to in his rebuttal: without a fixed definition, the claim can't be checked against anything, which makes it hard to call it a measurement rather than a marketing line[9].
There is also a narrower, checkable dispute buried in OpenAI's own numbers. On several individual benchmarks, Astra leads outright. It scored 100% on ExploitBench, a test of finding and using software vulnerabilities, against 78.5% for OpenAI's prior model GPT-5.6 Sol and 70% for Anthropic's Claude Opus 5[8]. It also topped Terminal-Bench, OSWorld and FrontierMath, tests of operating a computer, using software tools, and solving advanced math[11].
But on the Artificial Analysis Intelligence Index, a broad average that blends performance across many kinds of tasks rather than just one, OpenAI's own comparison table shows Astra scoring 61.2 — behind both Claude Opus 5 and Claude Fable 5.1[16]. So the "who's ahead" question depends entirely on which chart you're looking at. Astra wins the tasks OpenAI's marketing leads with; Anthropic's models win the wider average. Both statements are true using the same underlying data.
The cybersecurity comparison has its own wrinkle. The Claude scores OpenAI used for its ExploitBench comparison came from Mythos, a restricted Anthropic build with fewer safety filters active, limited to a government-vetted cyber-verification program — not the public Fable 5.1 model most users would actually compare it to[16][17]. That distinction is disclosed, but only in a footnote, not in the headline chart[16].
A Model Built to Justify Its Own Price Tag
Astra was pretrained using more than 100,000 GPUs at OpenAI's Stargate site in Texas, with hundreds of thousands more Nvidia systems already planned[2]. Spending at that scale needs a continuous story of progress to justify it. A launch that looked routine would invite the question of whether the next round of that buildout is worth funding — so this one couldn't look routine.
Anthropic is under its own kind of pressure. The company confidentially filed paperwork for a stock market listing with the U.S. Securities and Exchange Commission on June 1, 2026, after a May funding round valued it at roughly $965 billion[14]. Reporting points to a listing as soon as October 2026, at a target valuation discussed as high as $2 trillion[14]. Anthropic's revenue run rate — sales projected out over a full year — climbed from about $9 billion at the end of 2025 to more than $65 billion by the end of July 2026[14]. In the weeks before that listing prices, every benchmark claim either company publishes is also a pitch to the bankers and investors who will set that number.
That helps explain why Anthropic shipped two days before OpenAI, and why OpenAI's launch leaned so heavily on the tasks where Astra wins. Both companies are self-grading against risk categories and benchmark suites that they, not any outside regulator, chose to publish. Sam Altman said Astra went through the White House's review process for advanced AI systems, but that process is voluntary — there is currently no U.S. agency that certifies a frontier model's capability or danger level before it ships[5].
What the "Critical" Label Actually Concedes
Security researchers and AI-risk critics read the Critical designation less as reassurance and more as an admission on the record. Their core point: OpenAI itself says Astra can find flaws nobody has published yet and build working exploits without step-by-step human guidance — and OpenAI shipped it anyway, restrictions or not[8]. A capability like that cuts both ways. It can help a company patch its own unknown vulnerabilities before someone else finds them, but it can just as easily help an attacker locate a target's weak point first[8].
Astra's own published system card adds a more technical concern: the model's reasoning has gotten harder to monitor for misalignment, meaning it is harder for OpenAI's own reviewers to check whether what the model says it's doing actually matches what it's doing[16]. Sanchit Vir Gogia of Greyhound Research offered a sharper framing of the Critical label itself: he argued it reflects a change in how OpenAI tested the model, not a jump in how dangerous the underlying model actually is between one assessment and the next[8].
Enterprise buyers, meanwhile, are watching a more practical number: cost per finished task, not leaderboard position. Astra reportedly matched Claude Fable 5 on a comparable coding benchmark while taking fewer steps to get there, and used fewer steps than both GPT-5.6 Sol and Claude Opus 5 on the same measure[16]. Fewer steps can mean a lower total bill even at a higher per-token price. Early users, though, report code that still needs cleanup, and a gap between the launch demo and daily use[9]. At $10 per million input tokens and $50 per million output tokens, a heavy workday of automated tasks adds up fast[11]. No independent study has yet measured what a model built for "computer use" — meaning it can browse, code, and operate software on its own — does to the jobs built around those same tasks[2].
How Each Newsroom Told the Same Story
Coverage split largely along the fault lines you'd expect, though not always in the way the labels predict. Fox Business led with performance gains and American industrial scale, treating the Critical classification as evidence of a careful company rather than a warning[6]. Breitbart put "AGI Era" in scare quotes and framed the whole story around Sam Altman's personal credibility, offering little of the benchmark detail a reader would need to judge the claim[10].
On the left, Gizmodo led with the AGI claim as a claim to be tested, giving Gary Marcus's rebuttal prominent space[9]. NBC News took a flatter approach, leading with the security trigger over the capability boast and correctly attributing the Critical label to OpenAI's own classification rather than stating it as settled fact[5]. Al Jazeera's headline paired the launch with "rising scrutiny and safety concerns," framing the release inside a governance problem rather than a product story[7]. CNBC and trade outlet CSO Online scored as the most measured of the coverage — CSO Online, in fact, is where the sharpest technical skepticism showed up, in the analyst's point that the Critical label may reflect changed testing rather than a changed model[4][8].
What's Still Unmeasured
A week after launch, the two companies' benchmark claims remain unverified by anyone outside the companies that produced them. No outside body checked Astra's exploit scores, and none is required to before a model like this reaches paying customers[8]. Anthropic's IPO filing is still confidential, and its October listing target is not locked in[14]. Whether Astra's agentic gains translate into real job displacement, and whether its harder-to-monitor reasoning becomes a practical problem rather than a footnote in a system card, are both questions with no data yet — just two companies, mid-fundraise, telling their strongest version of the same numbers.
Summary
OpenAI released a new flagship model, GPT-6 Astra, on September 3, 2026[1][2]. It first went to a limited set of trusted partners. Paying ChatGPT tiers — Plus, Pro, Business and Enterprise — plus the API and Amazon Web Services got access over the following days[2]. The model has a 1-million-token context window, a knowledge cutoff of April 30, 2026, and costs $10 per million input tokens and $50 per million output tokens[11]. A token is roughly three-quarters of an English word, so a million tokens is about 750,000 words — a long novel. At those prices, feeding in a novel-length document costs about $10, and getting a novel-length answer back costs about $50.
OpenAI says Astra is the first model to cross its own "Critical" cybersecurity threshold[8][13]. That label comes from OpenAI's internal risk policy, not from any government. It means the company's testers found the model could locate software flaws nobody had published yet and work out how to exploit them, without a human walking it through each step[8]. OpenAI says it therefore restricted access to the most offensive capabilities and shipped a version that refuses some cyber prompts[2]. Sam Altman said the model went through the White House's voluntary review process for frontier AI systems[5].
The loudest dispute is over what the benchmark numbers mean. OpenAI President Greg Brockman said "it's not unreasonable to feel that we are now in the AGI era," and Nvidia CEO Jensen Huang posted "AGI has arrived" on September 6[3][9]. Critics including Gary Marcus reply that there is no agreed definition of AGI, so the claim cannot be checked[9]. There is also a narrower, checkable dispute: Astra tops several reasoning and agent benchmarks, but on the Artificial Analysis Intelligence Index — a cross-domain average — OpenAI's own comparison table places Astra behind Anthropic's Claude Opus 5 and Claude Fable 5.1[16].
The timing matters commercially. Anthropic launched Claude Fable 5.1 and Mythos 5.1 on September 1, two days earlier[17]. Anthropic confidentially filed for an IPO on June 1, 2026, and reporting points to an October listing[14]. Every leaderboard claim in this window is also a pitch to investors.
The Event
OpenAI released GPT-6 Astra on September 3, 2026, first as a limited preview for trusted partners, then to paying ChatGPT users the next day in a restricted version that refuses some cybersecurity prompts[2]. Access through the OpenAI API, Amazon Web Services and the ChatGPT Plus, Pro, Business and Enterprise plans followed over the next several days[2]. OpenAI said the model was pretrained on more than 100,000 GPUs at its Stargate site in Texas and is the first model to cross the company's internal "Critical" cybersecurity risk threshold[2][8]. OpenAI President Greg Brockman said it is "not unreasonable to feel that we are now in the AGI era"[9].
Undisputed Facts
- GPT-6 Astra was released on September 3, 2026, initially as a limited preview for trusted partners[2].
- API pricing is $10 per million input tokens and $50 per million output tokens; the context window is 1 million tokens and the knowledge cutoff is April 30, 2026[11].
- OpenAI states that Astra is the first of its models to cross the company's internal "Critical" cybersecurity threshold, and that it restricted access to the most powerful offensive capabilities as a result[8][13].
- In testing without production safeguards, OpenAI reported Astra scored 100% on ExploitBench, versus 78.5% for GPT-5.6 Sol and 70% for Claude Opus 5, and 42.4% on ExploitGym versus 30.3% for Sol[8].
- OpenAI reported Astra scored 57.9% on Terminal-Bench 4.0 (Sol: 37.3%), 72.6% on OSWorld 2.0 (Sol: 65.7%) and 97.6% on FrontierMath Tier 4 (Sol: 83.0%)[11].
- On the Artificial Analysis Intelligence Index, a cross-domain average, OpenAI's own comparison table places Astra at 61.2, behind Claude Opus 5 and Claude Fable 5.1[16].
- Anthropic announced Claude Fable 5.1 and Mythos 5.1 on September 1, 2026, two days before the Astra launch[17].
- Anthropic confidentially filed an S-1 with the U.S. Securities and Exchange Commission on June 1, 2026, after a May 2026 round set a post-money valuation of roughly $965 billion[14].
- Sam Altman said Astra went through the White House's voluntary vetting process for frontier AI systems[5].
The Pressure
Strip away the moralizing and blame. What structural realities persist regardless of which narrative wins?
- Capital has to be justified
- Astra's pretraining ran on more than 100,000 GPUs at Stargate in Texas, with hundreds of thousands more Nvidia systems planned[2]. Spending at that scale requires a continuous story of progress. A launch that looked incremental would raise the question of whether the next buildout is worth funding — so the launch cannot look incremental.
- The IPO window
- Anthropic filed confidentially on June 1, 2026, with an October listing expected[14]. In the weeks before pricing, a competitor's benchmark claims are not just technical — they are inputs to how bankers and buyers value the company. That is why Anthropic shipped on September 1 and OpenAI on September 3.
- Self-assessment is the only regime that exists
- The "Critical" threshold is OpenAI's own category, measured by OpenAI's own tests[13]. The White House review Altman cited is voluntary[5]. No U.S. agency currently certifies frontier model capability or risk. So the same company both sets the bar and reports whether it was crossed.
- Benchmarks are chosen, not neutral
- Which benchmark you lead with decides who wins. Astra tops agent and math benchmarks; Anthropic's models top the broader Artificial Analysis average[11][16]. Both statements are true at once, and each company's press materials feature the one it wins.
- Offense-defense asymmetry in cyber
- A model that finds unpublished software flaws helps both defenders patching their own code and attackers hunting targets[8]. The same capability serves both. Restricting the offensive version limits who gets it — but it does not change what the underlying model can do.
Material realityA model exists that OpenAI's own testing says can find previously unknown software vulnerabilities and devise exploits without step-by-step human direction[8]. That is on the record from the vendor, whatever the AGI debate resolves to. Its measured performance is genuinely mixed: leading on agent, coding and math benchmarks, trailing two Anthropic models on a broad cross-domain average, per OpenAI's own table[11][16]. Access is priced at $10 per million input and $50 per million output tokens through the API, AWS and paid ChatGPT tiers[2][11]. Two of the three companies most affected — OpenAI and Anthropic — are in a capital-raising cycle where these numbers carry billions of dollars of valuation. No independent body verified any of the benchmark claims before publication, and none is required to.
Narrative as a weaponThree parties are actively shaping this. OpenAI wants you to believe a capability line was crossed and that it crossed it responsibly — the AGI framing and the voluntary "Critical" disclosure are the same message from two directions. Nvidia's Jensen Huang, whose hardware trained the model, amplified the strongest possible version of that claim with "AGI has arrived"; his company sells the shovels either way. Anthropic wants you to look at the broad average rather than the headline benchmarks, and to notice that the cyber comparison used its restricted Mythos build rather than the public model — a correct point that also happens to be the one that protects its IPO. AI-risk critics want you to read the disclosure as an admission rather than a credential. Their sharpest evidence is not the AGI dispute at all but the system card's own note that the model's reasoning got harder to monitor. Ordinary readers should hold two things at once: the benchmark gains are real and documented, and every party publishing those numbers has money riding on how you read them.
How Each Side Sees It
Each major actor’s view — how it frames things, its underlying incentive, and how it’s materially affected. Tap a side to read it.
Frames it asOpenAI argues that capability and safety advanced together here. It says it delayed the release after a July 2026 security incident to add safeguards, tested the model against its own published risk framework, self-declared the top "Critical" cyber tier rather than staying quiet, shipped a restricted public version, and reported two previously unknown vulnerabilities its own testing surfaced to the affected maintainers[2][8]. On capability, its case is that the tasks that matter are agentic — a model that operates a computer end to end, using apps and checking its own work — and that Astra leads exactly there, on Terminal-Bench, OSWorld and FrontierMath[11]. Brockman's AGI framing is offered as a description of what users will feel using it, not as a formal claim[9].
WhyOpenAI needs to justify enormous capital spending. Astra was pretrained on more than 100,000 GPUs at Stargate in Texas, with hundreds of thousands more Nvidia systems planned[2]. Setting the public narrative — and doing it two days after Anthropic's launch and weeks before Anthropic's expected IPO — protects both enterprise share and the fundraising story[14][15].
Impact on themA benchmark lead converts directly into enterprise deals and API volume at $10/$50 per million tokens[11]. The "Critical" self-designation is a two-sided bet: it builds credibility with regulators, but it also puts on the record that OpenAI shipped a model it says can find and exploit unknown flaws[8].
Frames it asAnthropic's strongest argument is that the comparison is not apples to apples. On the Artificial Analysis Intelligence Index — a broad average across domains rather than a single task — Claude Opus 5 and Fable 5.1 sit ahead of Astra, by OpenAI's own table[16]. On the cyber comparisons, the Claude scores OpenAI used came from Mythos, a build running with fewer safety classifiers active and restricted to the U.S. Cyber Verification Program, not the public Fable 5.1 — a distinction disclosed in a footnote, not the headline table[16][17]. Anthropic's wider position is that gating dangerous capability behind verification programs is the responsible design, and that being outscored on an unsafeguarded exploit benchmark is a consequence of that choice, not a failure.
WhyAnthropic filed confidentially for an IPO on June 1, 2026, with reporting pointing to an October listing[14]. Its pitch rests on enterprise trust and safety differentiation. A rival's headline benchmark sweep in the pricing window is a direct threat to that story.
Impact on themReported annualized revenue passed $65 billion by end of July 2026, up from about $9 billion at the end of 2025[14]. A perception that OpenAI has retaken the frontier can move the valuation range — the gap between the roughly $965 billion May mark and the $2 trillion figure discussed for the listing is exactly the space this narrative fight plays in[14].
Frames it asThis camp says the disclosure proves the problem rather than solving it. The point they press hardest: OpenAI itself says the model can find unpublished flaws and exploit them without step-by-step human direction, and OpenAI shipped it anyway[8]. Astra's own system card, they note, documents a decline in how easily the model's reasoning can be monitored for misalignment — meaning it is getting harder to check whether the model's stated reasoning matches what it is actually doing[16]. Sanchit Vir Gogia of Greyhound Research adds a sharper point: the "Critical" label was a disclosure event, not a capability change — the underlying model did not get more dangerous between assessments; the testing method changed and produced a reclassification[8]. On AGI, Gary Marcus's argument is that there is no agreed definition, so the claim can be neither verified nor falsified — which makes it marketing, not measurement[9].
WhyTo force external, independent evaluation instead of company self-assessment, and to slow a release cadence they say is outrunning any ability to check it[7].
Impact on themTheir leverage is regulatory. The White House process Altman cited is voluntary[5]. If a real-world incident is traced to a frontier model's cyber capability, this group's framing becomes the basis for binding rules.
Frames it asBuyers care about cost per completed task, not leaderboard rank. The argument that matters to them is efficiency: Astra reportedly matched Claude Fable 5 on a comparable coding result while using fewer steps, and needed fewer steps than both GPT-5.6 Sol and Claude Opus 5[16]. Fewer steps can mean lower total cost even at a higher per-token price. Against that, early users report imperfect code and a gap between demo and daily use[9]. Worker-side arguments focus on the "computer use" pitch — a model that browses, codes, uses apps and runs long tasks is aimed at whole job functions, not single keystrokes[2].
WhyBuyers want to avoid locking into one vendor while the leaderboard flips every few weeks. Labor advocates want displacement effects measured before deployment, not after.
Impact on themAt $10 per million input and $50 per million output tokens, a heavy agentic workload runs into real money fast[11]. Displacement claims remain speculative — no independent study of Astra's labor effects exists a week after launch.
Like this article?
The Bias Ledger average rating 4
The same story, as framed by outlets across the spectrum, ordered least to most biased. The bias score (1 = straight, 10 = heavily spun) is an AI assessment of that framing — click an outlet to see its track record. The tell is the word choice or omission that reveals the angle.
| Outlet | Vantage | Bias | How they frame it | The tell |
|---|---|---|---|---|
| NBC News | U.S. center-left | 2 | "OpenAI debuts GPT-6 Astra, says it triggered security measures" — the safety trigger, not the capability, is the news. | "Says" correctly attributes the classification to OpenAI. Leading with security over performance is an editorial choice, but a defensible one, and the piece includes Altman's White House vetting statement. |
| CNBC | U.S. center, business | 2 | "OpenAI announces rollout of GPT-6 Astra model" — flat announcement framing with the cyber angle inside. | About as low-spin as the coverage gets. The market lens is the limit: competitive and valuation implications get more room than the safety debate. |
| CSO Online | U.S. trade press, security-practitioner audience | 3 | "OpenAI launches GPT-6 Astra, its first model to cross a critical cybersecurity threshold" — the threshold is the whole headline. | Audience-shaped emphasis, but it is the outlet that surfaced the strongest skeptical point: an analyst arguing the "Critical" label reflected changed testing methods, not a changed model. That specificity cuts against pure alarm. |
| Fox Business | U.S. right | 4 | "OpenAI rolls out GPT-6 Astra, touting major AI performance gains" — capability and American industrial scale lead. | "Touting" is a mild hedge, but the piece organizes around advances in coding, science and professional work. The "Critical" cyber classification appears as a feature of a careful company rather than as a warning. |
| Gizmodo | U.S. left | 5 | "OpenAI Claims We're in the 'AGI Era' With Release of GPT-6 Astra" — the story is the claim, not the model. | "Claims" plus scare quotes, then Gary Marcus as the rebuttal voice. The framing is defensible skepticism, but the model's actual measured gains get less space than the hype critique. |
| Al Jazeera | Qatari state-funded | 5 | "OpenAI unveils GPT-6 Astra amid rising scrutiny and safety concerns" — the launch is framed inside a governance problem. | "Amid rising scrutiny" is an editorial premise placed in the headline; the article's own quoted concern is that labs are not slowing down for cyber risk. Capability numbers are present but subordinate. |
| Fortune | U.S. center, business | 5 | "OpenAI launches GPT-6 Astra, its most powerful model yet, and touts its ability to use your computer" — carries OpenAI's superlative in the outlet's own voice. | "Its most powerful model yet" is stated flatly rather than attributed, and Brockman's AGI line is given prominence. The Artificial Analysis result placing Astra behind two Anthropic models is not the framing. |
| Breitbart | U.S. right (populist) | 6 | "Sam Altman's OpenAI Releases 'Astra' AI Model Claiming the 'AGI Era' Is Here" — the claim is attributed to a named man and quarantined in quotation marks. | Personalizing the company as "Sam Altman's OpenAI" and scare-quoting "AGI Era" frames the launch as an elite credibility question. The benchmark detail that would let a reader check the claim is thin. |
References
- GPT-6 Astra: A new generation of intelligence — OpenAI · primary source — the company launching the product
- GPT-6 Astra — Wikipedia · crowd-edited encyclopedia; aggregates cited reporting, not a primary source
- OpenAI launches GPT-6 Astra, its most powerful model yet, and touts its ability to use your computer — Fortune · U.S. business press, subscription and advertising funded
- OpenAI announces rollout of GPT-6 Astra model — CNBC · U.S. business news, owned by Comcast/NBCUniversal
- OpenAI debuts GPT-6 Astra, says it triggered security measures — NBC News · U.S. center-left broadcast news, owned by Comcast/NBCUniversal
- OpenAI rolls out GPT-6 Astra, touting major AI performance gains — Fox Business · U.S. right-leaning business network, Fox Corporation
- OpenAI unveils GPT-6 Astra amid rising scrutiny and safety concerns — Al Jazeera · Qatari state-funded international broadcaster
- OpenAI launches GPT-6 Astra, its first model to cross a critical cybersecurity threshold — CSO Online · U.S. IT security trade publication, advertising and vendor-marketing funded
- OpenAI Claims We're in the 'AGI Era' With Release of GPT-6 Astra — Gizmodo · U.S. left-leaning technology site, advertising funded
- Sam Altman's OpenAI Releases 'Astra' AI Model Claiming the 'AGI Era' Is Here — Breitbart · U.S. populist-right advocacy outlet
- GPT-6 Astra Benchmarks Explained — Vellum · commercial AI tooling vendor; publishes benchmark write-ups as marketing content
- OpenAI Launches GPT-6 Astra After A Curious False Start — Forbes · U.S. business magazine; contributor network with variable editorial control
- Safety overview: GPT-6 Astra — OpenAI · primary source — the company's own safety self-assessment
- Anthropic IPO 2026: Plans September or Early October Listing Amid $965 Billion Valuation Talks — KuCoin · cryptocurrency exchange blog; commercial interest in trading interest, not an independent newsroom
- AINews: GPT-6 Astra — OpenAI's biggest LLM launch of all time — Latent Space · independent AI industry newsletter written by and for AI developers; enthusiast-leaning
- Anthropic GPT Race Splits the Benchmarks as Astra Resets the AGI Clock — Remio · AI product company blog; commercial content, analyzes published benchmark tables
- Claude Mythos — Wikipedia · crowd-edited encyclopedia; aggregates cited reporting, not a primary source