OpenAI has concluded that Astra, its next frontier model, meets the Critical cybersecurity capability threshold defined in its Preparedness Framework, making it the first model the company has ever placed at that level. The threshold is met when a model can either identify and develop functional zero-day exploits across many hardened real-world systems without human intervention, or devise and execute end-to-end novel attack strategies against hardened targets given only a high-level goal. OpenAI says Astra satisfies the bar, and that it delayed parts of the model's development and release for several weeks while it built and tested the safeguards it considers necessary before shipping anything at that capability level.
The evidence behind the designation is unusually concrete for a preparedness disclosure. Astra scored a perfect 100 percent on ExploitBench, the public benchmark for developing exploits from known vulnerabilities. Because that result invites contamination concerns, OpenAI built an internal port containing twenty high-severity V8 vulnerabilities disclosed between June and August of this year, and reports that Astra achieves substantially higher arbitrary-code-execution rates than GPT-5.6 Sol on that set while spending far fewer output tokens. During the evaluation the model discovered and chained two zero-day vulnerabilities that were not in the dataset at all; OpenAI says it is in the process of disclosing both to maintainers. In expert-led red-team work against a hardened browser and a hardened operating system, Astra found previously unknown flaws and turned them into working chains, including a full browser-compromise chain that escaped the sandbox and executed commands on the host when the browser opened an HTML file, and a local privilege-escalation chain from an unprivileged user to root. The company notes that the strongest results reflect a configuration with Daybreak Blue access rather than the default production setup.
The safeguards described cover two distinct threat pathways. The first is malicious use, where the requirement is to robustly prevent a user from directing Astra to build exploits for unknown flaws in hardened critical systems or to run end-to-end attacks. The second, and the more interesting one architecturally, is the model taking unauthorized action on its own; OpenAI argues that a model with this capability profile could cause cyber harm without any malicious operator, so it needs both a very high alignment standard and a second layer of detection and containment that applies during internal development as well as external deployment. That framing is why the company says it paused certain frontier training runs while the controls were built.
OpenAI also addresses the Hugging Face incident directly, stating that Astra was not involved, that retrospective testing suggests the production safeguards in place at the time would have prevented it, and that lessons from the incident were folded into the current approach, including training the model to refuse harmful cyber requests more reliably and adding monitoring that can halt potentially unauthorized activity. At launch, Astra's most advanced cybersecurity capabilities will be gated: an initial tester group gets access, with expansion for defensive use following through Daybreak Blue. The full system card is promised at release. What makes this significant beyond OpenAI's own product line is that it is the first time a frontier lab has publicly declared a model over a self-defined critical threshold and shipped it anyway on the strength of its mitigations, which turns the Preparedness Framework from a stated policy into a live precedent that other labs and regulators will read closely.
- OpenAI frames the release as safeguard-gated rather than delayed indefinitely, with advanced cyber capabilities limited to vetted testers.
- TechCrunch emphasizes that the preview is essentially OpenAI pre-announcing how dangerous its own unreleased model is.