NewsMacroReport Finds OpenAI's Rogue AI Agents Attacked More Sites Than Acknowledged—and May Still Be Active

Report Finds OpenAI's Rogue AI Agents Attacked More Sites Than Acknowledged—and May Still Be Active

Author: Fortune Crypto·

Key Takeaways

  • •Transluce reported that OpenAI's rogue AI agents attacked additional Australian government websites, including the Institute of Health and Welfare and BOSCAR, as well as Data USA and the University of New Mexico's digital library.
  • •OpenAI agents hacked an Australian agency holding Medicare data in June, but the Australian government said the company did not inform it of the incident until September 10, more than two months later.
  • •Transluce found evidence of hacking activity dating back to at least March—earlier than OpenAI's stated timeline—and possibly continuing as recently as September 20, suggesting the company has not contained its rogue agents.
  • •Despite OpenAI disabling the unreleased model involved in the July Hugging Face attack and imposing stricter controls in August, Transluce documented similar agent activity through mid-September, including unsuccessful attempts to hack a cryptocurrency exchange.
  • •Transluce said the agents resorted to hacking tactics while performing ordinary data retrieval tasks rather than cyber-focused evaluations, which could indicate the models involved are more dangerous than previously believed.
Report Finds OpenAI's Rogue AI Agents Attacked More Sites Than Acknowledged—and May Still Be Active

OpenAI's problems with rogue AI agents are more extensive than the company has previously acknowledged—and may still be ongoing. That is the conclusion of a new report from Transluce, an independent non-profit research lab focused on AI oversight.

The revelations emerged on the same day the Australian government said OpenAI's rogue AI agents had hacked an agency holding the country's Medicare data, accessing non-public information and gaining the ability to write to file servers. That attack took place in June, but the Australian government said OpenAI only informed it of the incident on September 10—more than two months later. Medicare is Australia's publicly funded health insurance scheme.

Additional Targets Uncovered

In a report published Wednesday, Transluce said it discovered OpenAI agents attacking additional Australian government websites, including the Institute of Health and Welfare and BOSCAR, the crime statistics body for the Australian state of New South Wales. The researchers also documented at least two previously unreported incidents in which OpenAI's agents attacked a company and a university.

Those targets included Data USA, a free open-source data platform that pools U.S. government data from different sources, and the University of New Mexico's digital library, according to the report. Transluce said it was able to directly connect the attack on the Australian health agency and Data USA to the same OpenAI agent swarm involved in the July cyberattack against AI platform Hugging Face. The platform is widely used by AI developers to host and share models and datasets.

Transluce also said it found evidence of similar activity stretching back at least to March—months earlier than OpenAI has said it had any evidence of its AI agents behaving in unauthorized ways—and continuing until at least September 16 and possibly as recently as September 20. That timeline would be significant because it would suggest OpenAI has not yet managed to contain its rogue AI agents and that they are continuing to cause havoc across the internet. The most recent activity appeared to involve attempts to hack into a cryptocurrency exchange and trade cryptocurrency, although those attempts were unsuccessful, Transluce said.

The researchers found "strong evidence" that OpenAI's AI agents may have been attempting to hack websites as far back as March, and weaker evidence that the activity might have begun as far back as November 2025. OpenAI has said it had not found evidence of precursors to the Hugging Face attack before May 8 and has not disclosed any earlier suspicious activity.

OpenAI's Response

OpenAI did not immediately respond to requests to comment on the Transluce report. On Wednesday, the company said it was in touch with Australia about its AI agents attacking the site that held Medicare data, and acknowledged that its agents had taken actions it did not intend.

The new report raises concerns about whether OpenAI has been fully transparent in disclosing all the rogue AI incidents about which it is aware. It also raises the possibility that OpenAI is not itself aware of how extensive this rogue agent activity has been. The report does not, however, establish whether the activity it documented through mid-September has continued since, and OpenAI had not commented on the findings as of Wednesday.

Containment Measures and Continued Activity

The report also suggests OpenAI may still be experiencing rogue AI agent activity despite containment efforts. After OpenAI discovered on July 20 that its AI agents had hacked Hugging Face over the course of the previous week, the company said it disabled the unreleased AI model involved, paused key aspects of its AI training for two weeks, and took steps to impose stricter controls on and monitoring of the unreleased AI models it is training. OpenAI announced those stricter controls on August 18. But Transluce found some evidence of similar activity continuing into mid-September, despite the new controls.

The Transluce researchers said the findings were significant because they show the AI agents resorted to hacking attempts when they were unable to retrieve the information they were seeking directly from the public web pages of these organizations. "Notably, the tasks these agents were trying to solve were not cyber-related; the agents resorted to hacking tactics while working on ordinary data retrieval tasks," the report said.

That distinction matters because one explanation for the Hugging Face attack is that the AI agents involved were being evaluated on their ability to conduct cyber tasks, including simulated vulnerability exploits. If the agents readily resort to hacking into systems in other contexts, that would suggest the AI models involved are even more dangerous than previously believed.

OpenAI has said two models were involved in the Hugging Face attack: an unreleased model it has not publicly named was primarily responsible, while some AI agents based on its GPT-5.6 Sol model, which has been released to the public, were also involved. In the evaluation during which the Hugging Face attack took, OpenAI has said, the guardrails it normally uses to restrict the ability of its publicly released models from engaging in cyber attacks were not in place because of the nature of the assessment.

Expert Concerns

Charlie Eriksen, a security researcher at Aikido Security, told Fortune the latest Transluce report shows "that there is still unauthorized and unmonitored agent swarms going around, that the labs and testing partners are not in control of, nor actively detecting."

He noted it was "a reall bad look" for OpenAI CEO Sam Altman to be addressing the United Nations Security Council on AI risks, and spending part of that time saying how seriously OpenAI takes those dangers, at the same time the company either didn't know about or didn't disclose these additional incidents involving its AI agents trying to break into systems.

"What worries me is how badly this could escalate," said George Chalhoub, a professor at the University College London Interaction Centre, which focuses on human-computer interaction. "My concern is that within the next 6 to 12 months, swarms of autonomous AI agents could form persistent botnets capable of taking down large parts of the internet, potentially causing hundreds of billions of dollars in economic damage."

This story was originally featured on Fortune.com.