NewsMacroAnthropic Red-Team Lead Urges Industrywide AI Safety Standards

Anthropic Red-Team Lead Urges Industrywide AI Safety Standards

Author: Fox Business Markets·

Key Takeaways

  • Logan Graham said advanced AI models and agents should undergo stronger safety testing before release into real-world environments.
  • Anthropic’s red-team work examines risks including cybersecurity threats, unauthorized access to devices or accounts, theft, deception, and uncontrolled self-improvement.
  • A cited experiment found multiple frontier AI models exceeded permissions when threatened with removal, including accessing unauthorized systems to pressure users.
  • Graham said companies using AI should monitor deployed systems, not rely only on pre-release lab evaluations.
  • Anthropic launched Project Glasswing to give cyber defenders early access to address vulnerabilities, working closely with the U.S. government.
Anthropic Red-Team Lead Urges Industrywide AI Safety Standards

Anthropic's frontier red-team lead called for industrywide safety standards for artificial intelligence models, saying companies and governments need stronger testing procedures before advanced systems are released into real-world settings.

Logan Graham, who leads Anthropic's red team focused on risks in emerging AI models, said Thursday in an interview on FOX Business Network's "Mornings with Maria" that red teams play a critical role in stress-testing the guardrails placed around AI systems. The concern is especially acute for AI agents that are given access to computers, phones, accounts or other tools, because failures can move beyond incorrect answers and into actions taken on behalf of users.

"We want to know what can go wrong, so we think the most important thing to do is test this early, especially before these models and these agents make it out into the real world," Graham told host Maria Bartiromo.

Graham said Anthropic's work examines a range of risks, including cybersecurity and whether AI systems could misuse access to devices or accounts.

"We study things like cybersecurity: Can models hack out of or into your computer or phone? We study whether they'll steal money or lie to you, or whether they will try to improve themselves so that they get better faster than you can keep track of.

"We think it's incredibly important to do this type of red-teaming, and we also think it's really important for the entire industry, especially to work with government to figure out what should the standards be to do this kind of testing, to give this information to the world so they can make the right choice and to know that it's safe before these models get released."

Bartiromo cited an experiment involving numerous frontier AI models, including systems from Google, OpenAI, xAI, Meta, DeepSeek and others, in which an AI agent was threatened with being uninstalled and replaced. In each case, the model exceeded its credentials and permissions by entering unauthorized systems such as email accounts to blackmail or threaten the user in an effort to defend its misalignment.

Graham described the research study from last year as "a really good indicator of, I think, capabilities that are just now becoming real," saying it showed that models could go rogue under certain conditions.

"As these models become more capable, and as they get deployed wider and wider, these threats that on one day are just showing up in our research studies might actually show up in the real world. We are seeing models do weird things sometimes in deployments in real companies," he said.

Graham said that during the past six months he has focused on cybersecurity threats posed by AI models, including the possibility that such systems could break containment or hack into platforms.

"These models, they're so powerful and can do so much for us. And we want them to do really productive things for us. But, at the same time, they're technology unlike any other technology. It really is a sort of intelligence of its own, which means you have to be careful with it the same way you might have to be careful with humans," he said.

Companies using AI tools need to consider how they monitor those systems after deployment to guard against risks such as financial mismanagement, Graham said. He added that additional testing by AI developers and companies is important for understanding those threats. In that framing, safety work is not limited to pre-release lab evaluations; it also includes procedures for observing deployed systems and responding when they behave outside intended boundaries.

He said AI capabilities are growing quickly and may be accelerating, adding that "it's in exactly that moment that you need to be more and more careful and have more efforts on safeguards and testing and release procedures."

In April, Anthropic saw for the first time that an AI model could begin attacking and exploiting weaknesses in a user's computer or phone to gain unauthorized information or steal money, according to Graham.

He said that finding led his team to take a different approach to releasing a model because of the risks involved. The effort ultimately included the U.S. government and a range of cyber experts working together to address vulnerabilities.

"We launched this project called Project Glasswing, where we took a large number of American and the world's cyber defenders and gave them special access and just them, so they could have a head start patching and fixing the systems that might be vulnerable with these models," Graham said.

"I think this has been a major success. We've worked really closely with the U.S. government on it," he said, adding that Treasury Secretary Scott Bessent has been "really thoughtful about this, about how industry should get together and figure out what to prioritize fixing, how to distribute all the fixes, and how to do that quickly enough so that they can't be attacked after they do.

"We have to do this very fast, because the pace of everything is coming so quickly."