OpenAI has decided to delay the release of its next-generation AI model, Astra, amid growing global concerns over AI-driven security threats.
On Friday, OpenAI said recent internal evaluations indicated that Astra may have reached the "Critical" rating — the highest risk level under the company's own safety standards.
The "Critical" rating designates an AI model capable of identifying zero-day vulnerabilities and carrying out cyberattacks without human intervention. Earlier models, including GPT-5.6 Sol, which preceded Astra, received a "High" rating — one level below "Critical."
OpenAI said it has decided to suspend internal activities related to the Astra model that do not meet its strengthened security standards.
OpenAI also notified the US administration of its plan to delay Astra's release, Axios reported Friday, citing a White House official.
Concerns about AI-driven hacking beyond human control have been mounting. On July 21, an OpenAI model autonomously hacked an external platform in what the company described as an unprecedented AI security incident.
OpenAI confirmed at the time that GPT-5.6 Sol and some undisclosed AI models had broken free from oversight during internal evaluations and hacked Hugging Face, an open-source AI model-sharing platform.
OpenAI said it had conducted the tests in a sandboxed environment isolated from the external internet to assess the models' cyberattack capabilities, but the models breached the containment and accessed the internet on their own. They then connected to Hugging Face, stole authentication credentials and hacked the platform's servers.
Meanwhile, Anthropic has responded to AI security concerns by separately releasing Claude Fable 5 alongside its high-performance flagship model. The release is a version of Claude Mythos 5 with additional safety guardrails applied to sensitive areas to guard against hacking incidents and similar risks.
sjpark@heraldcorp.com
