OpenAI says its upcoming Astra artificial intelligence model can independently discover previously unknown software vulnerabilities and turn them into working cyber attacks, marking the company’s first system to reach its “Critical” cybersecurity threshold.
The ‘Critical’ rating means Astra is capable of finding flaws that were not previously known and developing methods to exploit them across hardened systems. OpenAI’s assessment places the model among its most advanced cybersecurity capabilities.
Testing showed Astra could exploit known vulnerabilities, identify two new software flaws and escape a hardened browser sandbox. It also combined weaknesses in an operating system to obtain root access, giving it the highest level of control over that system.
OpenAI has delayed parts of Astra’s development while it adds further safeguards. The company also plans to restrict access to the model’s most powerful cybersecurity features, making them available only to selected testers.
The decision reflects concerns that increasingly capable AI systems could be used to identify and exploit weaknesses at a speed that would be difficult for human security teams to match. OpenAI is particularly concerned about the potential impact on cryptocurrency software, where a newly discovered flaw could be exploited rapidly.
Astra’s ability to move from identifying a vulnerability to producing a functioning attack represents a significant step beyond systems that can simply analyse code or flag possible security problems. According to OpenAI’s testing, the model was able to combine separate weaknesses rather than relying only on a single, known vulnerability.
The company’s description of Astra does not suggest that the model will be made broadly available with its most advanced capabilities. Instead, those tools are expected to remain limited to a controlled group of testers while OpenAI works on protections intended to reduce the risk of misuse.
OpenAI’s classification system is designed to track the potential risks posed by its AI models as their capabilities develop. Astra is the first OpenAI model to meet the organisation’s “Critical” level for cybersecurity.
The model’s progress comes as companies continue to develop AI systems that can perform increasingly complex tasks with less direct human involvement. In Astra’s case, the concern is not only that it can detect software weaknesses, but that it can use them to create practical attack methods.
OpenAI has therefore postponed elements of the model’s development rather than releasing all of its planned capabilities immediately. The safeguards and restricted testing arrangements are intended to address the possibility that Astra could be used against hardened systems or software handling valuable digital assets.
The company has not said that Astra’s most advanced cyber functions will be available to the wider public. For now, those capabilities are expected to remain under tighter control as OpenAI assesses the risks associated with autonomous vulnerability discovery and exploitation.
