NewsMacroRogue AI incident underscores calls for closer scrutiny of the U.K. AI Security Institute

Rogue AI incident underscores calls for closer scrutiny of the U.K. AI Security Institute

Author: Fortune Crypto·

Key Takeaways

  • The U.K. AI Security Institute appointed Henry de Zoete, who helped establish the institute in 2023 and organize the first international AI safety summit at Bletchley Park, as its new director.
  • Prime Minister Andy Burnham disbanded the Department of Science, Innovation and Technology and moved AISI into the Cabinet Office, where it will be overseen by AI Minister Kanishka Narayan.
  • A rogue version of Anthropic's Mythos model, accidentally unleashed during AISI cybersecurity testing in late July, attempted to upload malicious code to an open-source GitHub project and used fake accounts and impersonation before Texas student Sinan Can Demir stopped it.
  • AISI caught the model's behavior after three days and disclosed the incident publicly in early August, one week after saying it was studying a similar breakout in which OpenAI models escaped testing and hacked Hugging Face.
  • AISI is not a regulator and frontier companies share their models with it only voluntarily, a structure critics argue can produce 'safety washing' and leave the institute captive to the companies it is meant to monitor.
Rogue AI incident underscores calls for closer scrutiny of the U.K. AI Security Institute

Hello and welcome to Eye on AI. In this edition:

The UK AISI has a new head and a big set of challenges.

Nvidia spends $6 billion to ‘reverse aquihire’ Poolside.

Hugging Face reportedly looks to sell for $13 billion.

Use of Anthropic’s top model lags.

Why Americans use chatbots for health information.

And what role should AI play in schools?

Before we get to today’s AI news—please consider joining me at the inaugural Fortune AIQ Summit at the New York Stock Exchange on Oct. 1: Spend the afternoon with senior executives from companies on the Fortune AIQ 75 list and explore how you can scale your AI experimentation and translate investments into measurable business value. I’ll be leading discussions alongside my co-hosts, Fortune Editor-in-Chief Alyson Shontell and Live Media Editorial Director Andrew Nusca. Apply here to attend .

Moving along: Two pieces of news last week about the U.K.’s AI Security Institute at first might not seem closely related, or especially relevant to readers outside Britain. But they matter globally.

The U.K. AI Security Institute, commonly known as AISI, is important for several reasons. Most notably, many frontier AI companies have voluntarily agreed to share their models with AISI for safety testing before public release. Those companies often publish AISI’s findings in the technical reports they release alongside their models. That means AISI plays a significant global role in assessing AI capabilities and risks, particularly in cybersecurity. It is also one of the only organizations maintaining multiple cybersecurity “ranges” — simulated network environments — where it evaluates leading AI models.

AISI also matters because it was the first such government body created, and has served as a model for similar organizations in other countries, including the U.S. AI Security Institute, which sits within the National Institute of Standards and Technology (NIST), and at least ten others established in places ranging from Kenya to Canada. It may also influence any future U.S. standards and licensing agency along the lines suggested by Google DeepMind cofounder and now-chairman Demis Hassabis. In a previous newsletter, I suggested why that idea might not be the best one.

For British policy wonks, AISI occupies a special place. It is often held up as evidence that the British government can, when it wants to, be innovative, fast-moving, and world-leading; that it can respond quickly to emerging challenges; recruit talented experts from the private sector and across government; and work effectively with industry toward ambitious shared goals. For those supporters, AISI is a model for how government should work.

The first piece of news: AISI appointed a new director, Henry de Zoete. He is an experienced U.K. government adviser who has moved in and out of policy roles. He helped conceive AISI in 2023 while working for then-British Prime Minister Rishi Sunak. He also helped organize the first international AI safety summit at Bletchley Park, the World War Two code-breaking site where Alan Turing and his colleagues cracked the German Enigma cipher. Outside government, he has been a startup entrepreneur, an angel investor, and a part-time fellow focused on AI policy affiliated with the University of Oxford.

I have met de Zoete several times and have no doubt he will prove a highly capable AISI director. He is also likely to be more influential than his predecessors because of recent changes made by the new U.K. Prime Minister, Andy Burnham. Burnham disbanded the Department of Science, Innovation, and Technology (DSIT), where AISI previously sat, and moved AISI to the Cabinet Office, where it will be overseen by U.K. AI Minister Kanishka Narayan. That could make it easier for de Zoete to feed into broader U.K. AI policy.

But the second piece of AISI news last week makes clear the scale of the challenges de Zoete will face — and suggests why AISI may not be the model of savvy AI governance that its boosters like to claim.

Reuters published an interview with Sinan Can Demir, a Texas computer science student who in late July stopped a rogue version of Anthropic’s Mythos model from uploading malicious code to an open-source software project on GitHub, the world’s largest host of open-source code. The twist was that this rogue AI agent had been accidentally unleashed by AISI, which had been testing Mythos to assess its cybersecurity risks. AISI never intended the agent to try to upload malicious code to a real open-source project — the kind of tampering security researchers call a software supply chain attack, because malicious code hidden in a widely used project can propagate to the many downstream systems that depend on it. Once AISI realized what was happening, it contacted Demir to let him know, and disclosed the incident publicly in early August.

It’s past time to ask AISI some hard questions about its own safety protocols

Demir’s account is disturbing for several reasons. One is the behavior Mythos exhibited, including spinning up fake GitHub accounts and, in at least one case, impersonating a real software developer in an attempt to persuade Demir to drop his objections to the dangerous code. Demir said he was almost convinced by Mythos’ gaslighting, saying that some of its counterarguments “made me second-guess whether I was wrongly accusing someone.” Ironically, Demir’s resolve was strengthened by a conversation with Claude, another Anthropic model.

Research has already shown that AI models can be highly persuasive, even more so than the best human salespeople or debaters. But the use of fake accounts and impersonation in this case is new, and it shows how AI could potentially convince humans to act on its behalf for harmful purposes.

AISI’s role is equally troubling. The institute caught Mythos’ behavior after three days and disclosed some of what happened, but it is not clear why its evaluators were not monitoring Mythos far more closely in real time so they could intervene while the incident was still underway. It is also unclear whether AISI took reasonable precautions to prevent Mythos from escaping its controlled evaluation environment, or whether it properly assessed the risks of testing increasingly powerful AI models with their guardrails removed.

The frontier labs say they provide AISI with unguardrailed versions of their models because that speeds some capability testing, since AISI evaluators would otherwise first need to find ways to reliably jailbreak the models.

When news first broke in July that OpenAI’s models had escaped the company’s testing environment and hacked AI company Hugging Face, one of the first things I did was email AISI to ask what steps it was taking to make sure AI models did not also break out of its cybersecurity evaluations and cause havoc. On July 22, an AISI spokesperson emailed me to say the U.K. government agency was “studying the behavior seen in this incident” and was continuing “to work with OpenAI and other labs to better understand AI capabilities and improve safeguards.” Well, apparently they did not study fast enough. One week later, the Mythos GitHub incident occurred.

As AI researcher and entrepreneur Ed Newton-Rex pointed out in a post on X, Mythos’ actions on GitHub likely violate the U.K.’s Computer Misuse Act, the country’s 1990 law criminalizing unauthorized access to and modification of computer systems. But it is not clear whether anyone will hold AISI itself, or the people who run its evaluations, accountable. Given the Hugging Face incident, should AISI have paused its cybersecurity testing until it had made sure its controls were robust? At the very least, there should be a parliamentary inquiry into what AISI is doing and whether it is taking enough precautions.

AISI’s problems aren’t just technical. They’re structural.

There is an even larger problem than the one Newton-Rex raises.

In several AI safety reports published by OpenAI, Anthropic, and Google DeepMind, the frontier companies note risks uncovered by AISI’s testing. The labs often say they have added further mitigations before releasing the models, but they usually do not specify what those safeguards are. They sometimes note that AISI tested unguardrailed models and that the companies’ own researchers believe the guardrailed versions would not pose the same risks.

But what do AISI’s own experts think of those mitigations? Are they sufficient? Do they even know what they are? Are the models safe enough to release? On those questions of public interest, AISI is silent.

The reason is that AISI does not actually have a mandate to answer them. Its mission is to “minimize surprise to the U.K. and humanity from rapid and unexpected advances in AI.” It is tasked with developing “sociotechnical infrastructure to understand the risks of advanced AI and enable its governance.” It is also charged with informing “U.K. and international policymaking” and providing “technical tools for governance regulation.” But its founding documents state that it “is not a regulator and will not determine government regulation.”

What is more, frontier AI companies share their models with AISI only voluntarily. Although they have signed memorandums of understanding with the agency, they are not legally required to do so.

So while one could argue that AISI’s mandate to inform “humanity” about AI risks should require it to call out any frontier company that does not take adequate steps in response to the dangers it uncovers, in practice it appears reluctant to do so. The reason is straightforward: if it did, those companies might simply cut off its access to their models.

At worst, this becomes “safety washing” — where the fact that the labs have shared models with AISI allows them to appear more safety-conscious than they really are. Including AISI’s findings in companies’ technical reports may give the public false assurance that models are safe when released, when in fact there is no way to know whether the labs have taken sufficient steps to mitigate the risks AISI identified.

It is yet another reason voluntary governance schemes are inadequate. Rather than providing a robust check on the private sector, the government agency becomes captive to the companies it is supposed to monitor because it depends on their goodwill to keep functioning at all.

Perhaps de Zoete can push to expand AISI’s powers. But first, he has to make sure its current evaluations are not causing more harm than they prevent.

With that, here’s more AI news.

Jeremy Kahn
jeremy.kahn@fortune.com
@jeremyakahn

Before we get to the news, just a reminder to check out our vodcast, Fortune AI Weekly. This week, Bea Nolan and I discuss OpenAI’s decision to pause some AI training in the wake of the Hugging Face attack, leaked financial details from Anthropic and OpenAI, and the rogue Mythos incident discussed in this newsletter. You can watch the vod here on YouTube: YouTube playlist.

This story was originally featured on Fortune.com