NewsMacroOpenAI Paused Certain Astra Activities Over 'Critical' Cybersecurity Findings; Training Continued, Altman Says

OpenAI Paused Certain Astra Activities Over 'Critical' Cybersecurity Findings; Training Continued, Altman Says

Author: CryptoBriefing·

Key Takeaways

  • OpenAI said Astra’s cybersecurity abilities reached or approached the "critical" level in its Preparedness Framework.
  • The model showed advanced agentic coding skills and the potential to find or exploit zero-day vulnerabilities.
  • OpenAI paused the affected internal activities for a little more than two weeks while it implemented new security measures.
  • Sam Altman said Astra’s core training did not stop and that new models are still expected to ship soon.
OpenAI Paused Certain Astra Activities Over 'Critical' Cybersecurity Findings; Training Continued, Altman Says

OpenAI halted certain internal activities connected to its Astra model on August 7, 2026, after preliminary findings indicated the AI had reached what the company classifies as "critical" cybersecurity capabilities. CEO Sam Altman, however, emphasized that Astra's core training never stopped and that new models remain on track to ship soon.

What triggered the pause

The concern centers on Astra's cybersecurity capabilities, which reportedly reached or approached the "critical" tier in OpenAI's Preparedness Framework — the highest classification the company uses to evaluate model risks.

In practical terms, the model demonstrated sophisticated agentic coding skills and the potential to identify or exploit zero-day vulnerabilities. Zero-day vulnerabilities are security flaws that software vendors don't yet know about, making them extraordinarily valuable to both defenders and attackers. AI-driven vulnerability discovery is already an active field of work: DARPA's AI Cyber Challenge, whose finals were held at the DEF CON hacking conference in 2025, tasked AI systems with finding and patching flaws in widely used open-source software, underscoring both the defensive promise and the dual-use risk of the capability at issue here.

The pause lasted slightly more than two weeks while OpenAI implemented new security measures around those specific capabilities. During that window, the company assessed the risks and put guardrails in place before resuming the affected activities.

Astra's broader capabilities

The cybersecurity dimension is only one facet of what appears to be an exceptionally powerful model. An earlier version of Astra reportedly solved 10 major open problems in mathematics and theoretical computer science — a claim that, if independently verified, would represent a landmark achievement in AI research.

Altman has said Astra is intended to be made "generally available," meaning it will not be locked behind restricted research access indefinitely. He acknowledged, though, that more time is needed to ensure the model is developed safely, with a particular focus on addressing its cyber capabilities before a broader rollout.

The safety-speed balancing act

OpenAI's Preparedness Framework was designed for exactly this kind of situation. Introduced in October 2023, the framework establishes risk tiers, from "low" to "critical," across categories including cybersecurity, persuasion, autonomy, and biological threats. When a model reaches the critical tier in any category, additional review and mitigation steps are triggered before deployment can proceed.

Comparable schemes now exist across the frontier-lab landscape: Anthropic's Responsible Scaling Policy and Google DeepMind's Frontier Safety Framework both define capability thresholds that trigger stricter safeguards, making tiered risk assessment an emerging industry norm. Astra's reported critical-tier finding is among the clearest public examples to date of such a threshold being invoked at a major lab.

The framework's prominence also follows an eventful period for OpenAI's safety governance. In 2024, the company disbanded its long-horizon Superalignment team after the departures of co-leads Ilya Sutskever and Jan Leike, folding its work into broader safety efforts — a sequence that kept questions about the company's pace-versus-precaution balance squarely in public view.

What to watch from here

The near-term question is when Astra, or models derived from it, will actually ship to users. Altman's statement that new models are expected soon suggests the timeline has not shifted dramatically despite the pause.

Also worth tracking is whether OpenAI publishes detailed findings from its safety evaluation. A model that hits the critical tier in cybersecurity risk would appear to warrant a thorough public accounting of what was found and what mitigations were put in place — and, with Anthropic and Google DeepMind operating similar threshold-based frameworks, whether peer labs begin disclosing comparable evaluations of their own.