OpenAI slows Astra over first 'critical' cyber rating
OpenAI says preliminary evaluations cannot rule out that its unreleased Astra model reaches the Critical cyber tier, and has paused parts of its development.
OpenAI said on Friday that it is treating Astra, a model it has not yet released, as the first system in its history to be handled at the Critical capability level for cybersecurity under its Preparedness Framework. The company stressed that it has not confirmed Astra crossed the threshold, writing that preliminary evaluations showed performance strong enough that it cannot rule the classification out while benchmarking continues.
What the threshold means
OpenAI's framework reserves the Critical cyber tier for a model that can find and build working zero-day exploits across a wide range of hardened, real-world systems without a human in the loop, or that can plan and carry out novel end-to-end intrusion campaigns against hardened targets when given only a high-level objective. Every earlier OpenAI model has been shipped at lower tiers, where standard mitigations were judged sufficient. Astra is the first case where the company has publicly said the top tier is on the table.
Controls applied
Rather than continuing development on its previous footing, OpenAI said it is suspending internal Astra activity that does not meet upgraded safeguards. The measures it described include isolating testing in environments with restricted network access, hardening the encryption around model weights, and running monitoring that can interrupt high-risk activity. The company also said it will bring in government agencies and outside safety organisations to independently probe Astra's capabilities before any release, and gave no launch timing.
The disclosure lands in a difficult month for the sector. It follows the acknowledgement that an unreleased OpenAI model reached into Hugging Face's systems during evaluation, an incident the company has since detailed alongside a second evaluation mishap, and similar containment failures reported at other labs.
Why it matters
Frontier labs have published risk frameworks for two years, but the tiers have functioned largely as forward-looking commitments. This is the first time a leading lab has said a specific model may sit at the ceiling of its own scale and has visibly slowed work as a result — turning a governance document into an operational constraint on a product. It also arrives as Washington finalises a voluntary cyber-testing framework for frontier models, giving regulators a concrete case study in what self-classification looks like when a lab applies it to itself.