Last Updated on September 13, 2026
OpenAI’s GPT-6 Astra has crossed a cybersecurity threshold that the company had previously reserved for a more serious level of AI capability. OpenAI classifies Astra as Critical for cybersecurity under its Preparedness Framework, meaning the model can, with suitable tools and access, find previously unknown vulnerabilities and develop ways to exploit them across hardened systems without a person directing every step.
That does not mean GPT-6 Astra is independently certified as a dangerous cyber weapon, nor does it mean the public version of ChatGPT can freely conduct such attacks. The Critical label comes from OpenAI’s own risk framework and its evaluations. The distinction matters because the evidence is based largely on OpenAI’s testing, although the company also used third-party assessments.
The development is significant for both attackers and defenders. A model that can find new software weaknesses can potentially help security teams discover and fix flaws faster. The same capability could reduce the expertise and time needed to develop attacks.
What does “Critical” mean for GPT-6 Astra?
OpenAI’s Preparedness Framework uses capability thresholds to assess models that could create severe risks if misused. For cybersecurity, its Critical threshold covers models that can identify and develop functional zero-day exploits across many hardened real-world critical systems without human intervention, or devise and execute new end-to-end attack strategies against hardened targets.
A zero-day is a previously unknown software vulnerability, or a flaw for which a fix may not yet be available. Finding one is difficult even for skilled security researchers.
OpenAI says Astra reached this threshold during its evaluations. That is different from saying the model has successfully carried out a real-world cyberattack. The testing was performed under controlled conditions, and OpenAI has restricted some of Astra’s more advanced offensive capabilities in its public products.
Astra’s cybersecurity results go beyond a benchmark score
OpenAI reports that Astra scored 100% on ExploitBench, compared with 78.5% for GPT-5.6 Sol. On the harder ExploitGym evaluation, Astra achieved 42.4%, compared with 30.3% for GPT-5.6 Sol.
Those figures need context.
OpenAI itself says that some older vulnerability benchmarks can be affected by a model having seen information about historical vulnerabilities during training. To test whether Astra could generalize to newer vulnerabilities, the company created an internal evaluation using vulnerabilities disclosed between June and August 2026.
During that testing, Astra discovered and used two previously unknown zero-day vulnerabilities as part of its exploit chains. OpenAI says it is disclosing those vulnerabilities to the relevant maintainers.
That is arguably more informative than the 100% benchmark result. It tests whether the model can work with vulnerabilities that were not part of the older information available to it.
Still, these results have not been independently replicated at the same scale. They should therefore be treated as OpenAI’s reported evaluation results, not as an independent measurement of Astra’s real-world attack capability.
Why this matters to cybersecurity teams
The defensive side of the story is easy to miss.
OpenAI says Astra can help with tasks such as secure code review and patching. Its ability to identify weaknesses can give defenders another way to find problems before attackers do.
But the same capability changes the risk calculation.
If AI systems become better at discovering vulnerabilities, attackers may eventually need less specialist knowledge to find and exploit weaknesses. Security teams may also face a shorter window between the discovery of a vulnerability and attempts to exploit it.
That makes basic security practices more important, not less. Organizations using AI agents for coding or security work need to think carefully about what systems those agents can access, what credentials they receive, and which actions require human approval.
OpenAI has restricted Astra’s strongest cyber capabilities
GPT-6 Astra is not being released with unrestricted offensive cybersecurity functionality.
OpenAI says Astra currently refuses advanced requests such as creating proof-of-concept exploits. The company plans to expand defensive cybersecurity capabilities through its Daybreak program, including vulnerability validation, malware analysis and detection engineering.
OpenAI has also added stronger monitoring and controls. Its safety documentation describes tighter isolation, encrypted model checkpoints, monitoring of tool-using Astra activity and additional evaluations designed to detect misaligned behavior.
There is another concern buried in the safety documentation. OpenAI says Astra’s written reasoning was harder to monitor than GPT-5.6 Sol’s in certain tests designed to evaluate attempts to evade monitoring. The company says improving monitorability remains an active research priority.
That matters because safeguards are only useful if organizations can reliably detect when an AI system is moving outside its intended scope.
The latest development: EU cybersecurity officials are testing Astra
The story has moved beyond OpenAI’s own testing.
On September 10, the European Commission said the European Union Agency for Cybersecurity, or ENISA, had been given access to GPT-6 Astra. ENISA is testing Astra alongside Anthropic’s Mythos 5 to assess their cybersecurity capabilities and possible impact.
That independent government testing could become important as more information emerges.
The strongest evidence available today still comes from OpenAI’s own technical documentation. External evaluations could provide a different view of how Astra performs under controlled security testing and whether the company’s reported capabilities translate consistently outside its own evaluation environment.
What GPT-6 Astra’s Critical rating does not prove
The announcement should not be read as proof that AI can now autonomously hack any system.
Astra’s capabilities depend on the tools, permissions, environment, and safeguards available to it. OpenAI’s own definition refers to appropriate access and tools.
The benchmark results also don’t establish that Astra will reliably perform the same tasks against every real-world target. Software environments vary, security controls differ, and successful exploitation depends on more than finding a theoretical vulnerability.
The better takeaway is narrower: frontier AI has reached a point where one major AI developer considers its cybersecurity capabilities capable of creating a new level of misuse risk.
That is already enough to change how AI agents should be deployed.
What should businesses watch next?
For organizations considering AI coding or cybersecurity agents, the most useful question is no longer simply whether a model is good at finding vulnerabilities.
The bigger question is what the model is allowed to do after it finds one.
Access controls, approval requirements, monitoring, and isolation become more important as models gain the ability to act across software systems. OpenAI’s Astra deployment shows why capability testing and security controls now need to develop together.
The next useful signal will be the results from external evaluations such as ENISA’s testing. Those findings can help determine how much of Astra’s reported capability holds up outside OpenAI’s own testing environment.
For now, GPT-6 Astra’s Critical classification is best understood as a warning about the direction of AI cybersecurity: models are becoming capable enough that the same system can become a valuable defensive tool and a serious security risk, depending on who controls its access.
- OpenAI Navier-Stokes Solution: What the AI Math Breakthrough - September 13, 2026
- The Current Z library Official Domain Has A New Website Address - September 12, 2026
- How Localization Technology is Encouraging the Global Outreach of Aviator - September 8, 2026


