共有:
OpenAI · Preparedness Framework

It Solved 10 Problems. Four Days Later,
OpenAI Called Its Own Model "Critical."

On August 3, OpenAI announced that its next flagship model, "Astra," had solved ten math and theoretical computer science problems that had stood unsolved for over a decade. Just four days later, before the excitement even faded, OpenAI reversed tone and said Astra had been flagged as a "critical cybersecurity risk" — the first time the company has ever applied that label to one of its own models.

AI Navigate Editorial·2026.08.11·6 min read
PREPAREDNESS FRAMEWORK — CYBERSECURITY LOW MEDIUM HIGH CRITICAL GPT-5.6-Sol Astra
01
Why It Matters

Why this is
a genuine first

A top-tier label, applied pre-emptively to a risk that hasn't happened yet.

In its Preparedness Framework, published in December 2023, OpenAI rates model risk across several domains — including cybersecurity — on a four-tier scale: Low, Medium, High, and Critical. Until now, the highest any model had reached in the cybersecurity domain was High, a level reached by OpenAI's recent flagship GPT-5.6-Sol. This time, OpenAI said in an official blog post that internal evaluations of its next model, "Astra," could not rule out the model reaching the Critical threshold. It's the first time OpenAI has acknowledged that possibility for any of its own models, as reported by GIGAZINE and other outlets.

The timing stands out. The announcement came just four days after OpenAI touted Astra's mathematical breakthroughs on August 3. Flaunting overwhelming capability one moment, then admitting a risk it can't fully contain the next — on paper, those two messages sit awkwardly together. It looks less like coincidence and more like a deliberate sequencing: impress the market with Astra's power, then immediately frame that same power as something being handled with extraordinary caution.

High (GPT-5.6-Sol)Critical (not ruled out for Astra)
Capability that amplifies existing attack methodsA capability that could open entirely new attack paths defenses aren't built for
Finding and executing zero-days still needs human involvementCould identify and execute zero-days without human involvement
Runs through normal development and deployment processRuns in isolated environments, with hardened weight encryption and monitoring

Not "we can still stop it,"
but "it may already be possible."


02
Timeline

What happened
in four days

From a math breakthrough to a cyber-risk warning — laid out in order.

2023.12 Framework published 2026.8.3 10 math results announced 2026.8.7 Cyber Critical flagged Since Development in isolated environments
FIG. From the math breakthrough announcement to the Critical designation, just four days apart

Since publishing the Preparedness Framework in December 2023, OpenAI has been rating its models against high-risk domains including cybersecurity, biological/chemical threats, persuasion, and self-improvement. Lined up in order, this run of announcements shows a shift in how the company is presenting its next model to the world.

On August 3, OpenAI published answers to ten long-standing open problems in mathematics and theoretical computer science, backed by a 249-page manuscript and Lean 4 formal proof certificates — including an explicit construction of a non-sofic group, resolving a question Mikhail Gromov posed about soficity back in 1999. It was the kind of result experts could verify line by line. Four days later, the same model's downside was the headline. Disclosing a model's strength and its risk this close together is, even by OpenAI's own standards, an unusually tight cadence.

03
By The Numbers

The first "Critical,"
by the numbers

1st ever
Critical designation for an OpenAI model
4 days
From math announcement to cyber warning
Top of 4 tiers
Low / Medium / High / Critical

OpenAI says that for capabilities judged Critical-level, it will pause parts of normal development in favor of stronger safeguards. The measures it describes include restricting access to isolated test environments, limiting network and tool access, hardening protection and encryption of model weights, and adding extra monitoring and detection. Internal activity that doesn't meet the stricter bar is reportedly being paused.

01

Isolated environments

Access to Astra is restricted to isolated test environments, cut off from the normal development pipeline.

02

Restricted access

Network and tool-execution access is trimmed down to the minimum necessary.

03

Stronger weight protection

Model weight encryption and protection are tightened to reduce the risk of exfiltration.

04

Heavier monitoring

Chain-of-thought monitoring and similar systems flag high-risk behavior; activity below the bar is paused.

04
Who's Affected

Who this actually
changes things for

Nothing changes in your everyday experience today. But the impact on different roles is not the same.

Enterprise risk teams

Procurement and security-review teams evaluating OpenAI now have an official risk classification they can cite in vendor risk assessments. For contract renewals and adoption sign-off, that's a practical, near-term win.

Product & business leaders

PMs planning around Astra or agentic-coding features should assume a more cautious release cadence than usual. Early access to coding- and automation-heavy capabilities in particular could be gated for longer than expected.

Everyday users

Nothing visible changes in day-to-day ChatGPT use right now. This designation is about how OpenAI manages development internally, not how the product behaves for casual users.


05
What To Do Next

What to do
about it now

1. Enterprise adoption teams should log this Critical designation and OpenAI's stated safeguards as an update to their vendor risk assessment and DPA review — it's likely to come up in the next security audit. 2. PMs and engineering leads planning around Astra should build slack into their roadmaps for GA and higher-tier features arriving later than expected. 3. It's worth continuing to watch for further Preparedness Framework documentation or statements from OpenAI's Safety Advisory Group, since more detail is likely to follow.

06
Caveats

This isn't a purely
reassuring story

A few caveats are worth keeping in mind. This designation comes from OpenAI's own internal evaluation; no independent third-party verification has been published as of this writing. "Cannot rule out reaching Critical" is a careful phrase — it signals an unconfirmed risk, not a confirmed one, and shouldn't be over-read as proof of an imminent threat. At the same time, it's worth not swinging all the way to "OpenAI erred on the side of caution, full stop," either. Announcing a headline-grabbing capability and then, days later, framing that same model as something being handled with extraordinary care is also a narrative that reinforces how far ahead Astra is — and that dual effect, intentional or not, deserves a measure of skepticism.