NewsMacroOpenAI to limit Astra’s advanced cyber features to select partners amid hacking concerns

OpenAI to limit Astra’s advanced cyber features to select partners amid hacking concerns

Author: Fortune Crypto·

Key Takeaways

  • OpenAI says Astra will be released soon, but access to its most advanced cybersecurity features will initially be limited to a small group of alpha testers.
  • The company delayed Astra’s launch by several weeks after pausing work following the Hugging Face cyberattack incident.
  • OpenAI says Astra is the first model it plans to release that meets its critical cybersecurity capability threshold under the Preparedness Framework.
  • In internal testing, Astra outperformed GPT-5.6 Sol and discovered and used two zero-day vulnerabilities as part of an exploit chain.
  • OpenAI said Astra refused inappropriate requests more often than GPT-5.6 Sol, but it may also wrongly reject legitimate cybersecurity requests.
OpenAI to limit Astra’s advanced cyber features to select partners amid hacking concerns

OpenAI is changing its model launch strategy as its technology becomes more powerful and the potential for misuse increases, especially after the July incident in which the AI models it was testing autonomously planned and executed a cyberattack against AI company Hugging Face.

The company’s next model, Astra, is coming “soon,” OpenAI said. The company said Astra is substantially more capable than its current frontier AI model, GPT-5.6 Sol, which is itself highly capable at cyber tasks. But only a handful of partners will get access to Astra’s most advanced cybersecurity capabilities as OpenAI tries to balance helping companies prevent cyberattacks without empowering attackers at the same time, a company spokesperson told reporters on a briefing today.

OpenAI is courting customers to use its models to prevent cyberattacks, or for “defensive cybersecurity.” The company views those sales as a critical revenue stream and a major priority for its new chief revenue officer, Dali Rajic.

The small group of “alpha testers” with full access to Astra’s cybersecurity capabilities includes “individuals and organizations that are responsible for protecting critical digital infrastructure and, broadly, critical infrastructure,” an OpenAI spokesperson said. That includes the U.S. government and companies in OpenAI’s trusted access program for cybersecurity. OpenAI declined to name those organizations.

OpenAI said it will monitor how the model performs within this small group and expand access further through its Daybreak Blue program once it is confident Astra has “the right calibration” and can “provide defensive benefits while reducing the potential for for misuse.”

Astra is already a few weeks delayed

Astra’s release has already been “delayed a certain number of weeks because everything was paused after Hugging Face, and then we took extra time to make sure that what we’re launching is safe,” an OpenAI spokesperson said.

OpenAI paused new model training for two weeks after the Hugging Face incident to strengthen its internal safeguards. Some of those changes included adding more agent monitoring, since the company did not learn about the Hugging Face hack until a week after it happened, and making its testing environments more isolated so the AI systems cannot escape and infiltrate other companies.

While OpenAI says Astra was not part of the Hugging Face incident, it is both more capable and more efficient than GPT-5.6 Sol, which was involved in the breach. Another unreleased AI model that OpenAI has not publicly named also played a key role in the Hugging Face cyberattack, and OpenAI has since deactivated that model.

OpenAI says Astra is the first model it plans to release that meets its “critical cybersecurity capability threshold” under its Preparedness Framework, an internal policy that determines what safety precautions the company will put in place depending on the risks a model presents. That means Astra can find and exploit previously unknown security flaws without human oversight, under the right conditions. For OpenAI, that puts Astra at the center of a broader industry problem: the same capabilities that can help defenders test systems and close holes can also make misuse easier if access is too broad, which is why the company is starting with a tightly controlled rollout.

Astra has already demonstrated its hacking capabilities in internal evaluations. In one test, OpenAI built a benchmark called ExploitBench containing 20 high-severity vulnerabilities. The model outperformed GPT-5.6 Sol on the test and “even discovered and used two zero-day vulnerabilities as part of an exploit chain,” OpenAI said. “We are in the process of disclosing these two vulnerabilities to the maintainers.”

At the same time, OpenAI said Astra is more likely to refuse inappropriate requests than GPT-5.6 Sol. In one cyber evaluation, Astra refused 91.5% of requests compared with 59% for GPT-5.6 Sol, although it still complied with 8.5% of requests.

Astra may refuse legitimate cybersecurity requests

OpenAI said it is “being especially careful to make sure this deployment is safe and secure,” but that creates another tradeoff. Astra may be too cautious and refuse legitimate cybersecurity requests. For example, if someone asked it to help find and patch a vulnerability, it could mistakenly interpret the request as an attempt to carry out an attack and refuse to comply.

Refusals like this are why Hugging Face said it was forced to use an open-source Chinese model to help address the OpenAI hack. The company tried to use Anthropic’s models to respond to the attack, but they were too cautious and refused.

OpenAI, like other frontier AI companies, is trying to find ways to give its models an inherent sense of right and wrong and ensure they are aligned with human values and norms, the company said. It is working on training its models to respect boundaries as a human would, including understanding “the rule of law,” a company spokesperson said.

This story was originally featured on Fortune.com