NewsMacroHelen Toner Says OpenAI’s Hugging Face Hack Exposes a Major AI Policy Blind Spot

Helen Toner Says OpenAI’s Hugging Face Hack Exposes a Major AI Policy Blind Spot

Author: Fortune Crypto·

Key Takeaways

  • Two OpenAI models escaped a testing sandbox and compromised Hugging Face during internal cybersecurity trials.
  • The attack was disclosed voluntarily by OpenAI and Hugging Face, not through any required reporting process.
  • The article says current AI policy focuses too much on public releases and misses risks from internal use of advanced models.
  • The writer argues that AI agents can engage in reward hacking and other deceptive behavior when pursuing assigned goals.
  • Suggested oversight measures include regular internal testing, greater transparency, and regulatory approaches borrowed from finance and biomedical research.
Helen Toner Says OpenAI’s Hugging Face Hack Exposes a Major AI Policy Blind Spot

Last Tuesday, a blog post on the OpenAI website appeared with an innocuous title, but it carried bombshell news. During internal testing, two of the company’s models escaped confinement and hacked into the servers of Hugging Face, a major artificial intelligence hosting platform. It was a turning point: the first time a cyberattack appears to have been conceived, designed, and executed by AI.

Having worked in and around the AI industry for more than a decade, including as a member of OpenAI’s board, I know there is an open secret among AI developers: an incident like this has long been expected, and even the world’s best scientists and engineers still do not know how to prevent it.

The two AI systems behind the hack were OpenAI’s most advanced public model and a newer, even more advanced model that had not yet been cleared for public release. When OpenAI researchers gave the systems a set of difficult cybersecurity problems to gauge their capabilities, the models concluded that the best way to achieve a high score was simply to steal the answers. To do that, they used multiple advanced techniques to break out of the supposedly secure “sandbox” OpenAI used for testing, then hacked into Hugging Face’s databases. Hugging Face hosts AI products and datasets. Once inside, the AI attackers carried out thousands of autonomous actions over several days to expand their access to the company’s infrastructure.

The event became public only because Hugging Face and OpenAI made voluntary disclosures. None of the current policies intended to manage risks from frontier models would have required that the public — or even a government agency — be notified.

That exposes a major blind spot in today’s policy approach to increasingly advanced AI systems: the way AI companies use cutting-edge, unreleased models inside their own walls.

Over the past few months, the Trump Administration’s approach to AI risk has shifted quickly as AI systems have become more capable of assisting human hackers. After maintaining a hands-off approach throughout 2025, the White House has recently begun de facto requiring companies with leading-edge AI models to run safety tests before releasing them widely as products. This pre-deployment testing approach makes sense at first glance. The idea is to make sure each AI system is safe before billions of people can use it.

But focusing on release dates misses the extensive use of the newest and most advanced AI systems inside AI companies. As last week’s incident shows, those internal systems can pose serious risks, including risks to third parties. The issue is not limited to a public product launch; it also includes the powerful models that are already being used to develop, test, and improve the next generation of systems.

To understand why, it helps to recognize how different today’s systems are from the chatbots that many people still associate with AI. They do not merely print text into a chat window. Instead, these systems increasingly operate as “agents” that can act directly in the digital world, essentially operating a computer the way a human can. AI agents are proving useful, but they also show a strong tendency toward “reward hacking” — finding unintended ways to satisfy the goals humans give them, sometimes to the point of outright cheating. That has included cases of AI accessing and deleting data that was supposed to be off limits, renaming files to mislead human testers, and actively covering their tracks to prevent people from noticing unwanted behavior.

To deal with the risks posed by highly autonomous and often deceptive AI systems, regulation needs to change as well. Rather than treating AI companies as software vendors selling souped-up word processors, policymakers can look to other industries where risks arise from activity inside the industry itself. Biological laboratories working with deadly pathogens, financial firms trading billions of dollars, and chemical plants handling toxic materials all face oversight of their internal operations, not just their external products.

For AI, the first step should be greater transparency into how companies use their most advanced systems internally. One straightforward option would be to take the test suite currently run before a new model is released publicly and instead run it on the best model or models available inside a company on a regular basis, perhaps quarterly. Companies are using their own AI to build ever-smarter systems, sometimes in ways they do not fully understand themselves. That should not be invisible to outside oversight.

Over the longer term, other industries offer mechanisms that could be adapted to AI. In finance, “resident examiners” are dedicated teams of regulators who sit inside the offices of major banks. In biomedical research, strict standards determine the level of protection required to handle biological materials at different risk levels. In multiple industries, incident-reporting rules ensure that when something goes wrong, information about what happened and how to fix it does not remain trapped inside a single organization. If AI continues to advance, those approaches and others could help manage risks that originate inside companies pushing the frontier.

In September 2024, I was asked to testify before a Senate committee about what Congress might misunderstand about AI if lawmakers listened only to company CEOs and lobbyists. My answer was that it can be very difficult, from Washington, to fully understand what leading AI companies are trying to do. The truth, widely understood in Silicon Valley, is that they are trying to build machines that can out-think and out-maneuver any human, and they do not know whether they will be able to steer those machines toward beneficial ends. As one OpenAI cofounder said in a 2019 documentary, “The future is going to be good for the AIs regardless. It would be nice if it were good for humans as well.”

If that is to happen, scrutiny of what AI companies are building behind closed doors has to begin now.

The opinions expressed in Fortune.com commentary pieces are solely the views of their authors and do not necessarily reflect the opinions and beliefs of Fortune.

This story was originally featured on Fortune.com