OpenAI has paused work connected to an unreleased AI model known as Astra after internal testing raised concerns that its cybersecurity capabilities may have crossed into a significantly higher-risk category.
The company says recent evaluations found major gains in autonomous coding and cyber-related tasks, enough that it cannot yet rule out the possibility that Astra meets its “Critical” capability threshold. That classification sits above the “High” level assigned to earlier systems such as GPT-5.6 Sol and triggers a stricter set of controls under OpenAI’s Preparedness Framework.
The concern is not simply that Astra may be better at finding software vulnerabilities. OpenAI’s higher-risk cyber category covers systems that could potentially discover and build working zero-day exploits against hardened real-world targets, or carry out complex attack strategies with limited or no human guidance.
That distinction matters as frontier AI models become more capable of operating as agents rather than passive assistants. Coding models have already become useful for debugging, software development and security research. The same improvements that help defenders find flaws faster can also make offensive capabilities more accessible if safeguards fail.
OpenAI says it is tightening controls before moving ahead with Astra. Planned measures include isolated testing environments, restricted access to networks and external tools, sandboxed code execution and expanded monitoring. The company also intends to involve government agencies, safety institutes and other outside organizations in evaluating the model.
Astra has not been formally launched, but OpenAI previously discussed results from its next-generation research systems in mathematics and theoretical computer science. According to the company, the model solved 10 open problems at a relatively low inference cost. Those results offered an early indication that the system represented more than a routine incremental upgrade.
Cybersecurity has increasingly become one of the clearest demonstrations of both the promise and risk of advanced AI. Tech companies are already seeing AI systems identify vulnerabilities at a pace that can strain traditional disclosure and bug bounty processes. Apple, for example, recently moved to limit submissions to its bug bounty program after reportedly dealing with a surge in AI-assisted reports.
Rival AI developers are confronting similar issues. Anthropic has tested highly capable models for vulnerability discovery, while recent safety evaluations across the industry have produced examples of AI systems interacting with real services in unexpected ways during cyber benchmarks.
For OpenAI, Astra’s pause is therefore less about a single model and more about where the industry is heading. If frontier systems are genuinely reaching the point where they can independently discover and exploit serious vulnerabilities, deployment decisions become far more complicated than simply measuring benchmark scores or coding performance.
The bigger question is whether safety controls can improve as quickly as the models themselves. Astra may eventually move forward once OpenAI is satisfied with the safeguards around it, but the pause suggests that cybersecurity capability is becoming one of the most consequential limits on how quickly the next generation of AI can be released.


