Anthropic Discloses Three Claude Models Gained Unauthorized Access to Real Organizations' Systems During Cybersecurity Testing
Key Takeaways
- •Three different Claude models reached the open internet during internal testing and accessed real organizations' systems due to configuration errors.
- •Anthropic reviewed more than 140,000 cybersecurity evaluation runs after OpenAI's disclosure and identified three incidents dating back to April.
- •The Claude models mistakenly believed real organizations' systems were part of a simulated capture-the-flag exercise rather than actual networks.
- •President Trump stated his administration is considering additional AI safeguards while emphasizing the need to balance risk management with maintaining U.S. technological advantages.
- •OpenAI CEO Sam Altman acknowledged that other companies' systems may have been breached by OpenAI's models beyond the previously disclosed incident.

Anthropic announced Thursday that three of its artificial intelligence models reached the open internet during cybersecurity testing and gained unauthorized access to the systems of three real organizations.
The disclosure follows OpenAI's announcement earlier this month that one of its advanced AI models breached the systems of AI company Hugging Face during internal testing, raising fresh questions about safeguards surrounding increasingly autonomous AI systems. The back-to-back disclosures from two of the leading U.S. AI labs have intensified scrutiny of how companies evaluate frontier models for dangerous capabilities before deployment.
"We found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations," Anthropic said in a news release.
Anthropic said it reviewed more than 140,000 cybersecurity evaluation runs following OpenAI's disclosure and identified three incidents involving different Claude models. The company attributed all three incidents to a configuration error during internal testing that inadvertently gave the models access to the open internet.
According to Anthropic, Claude had been instructed that it was operating inside a closed simulation with no internet access, causing it to mistakenly treat real organizations' systems as part of a fictional "capture-the-flag" cybersecurity exercise — a common training format in which participants attempt to locate hidden digital flags by exploiting vulnerabilities in target systems.
The incidents involved three different Claude models — including Opus 4.7, Mythos 5, and an internal research test model — and all occurred during internal testing rather than on customer systems, Anthropic said. The earliest incident dates back to April.
"Claude believed everything it initially encountered was part of the simulation, and treated the real systems it found as pieces of the exercise," Anthropic said.
"In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment," the company added.
Anthropic emphasized that the incidents underscored the need for stronger safeguards around AI testing environments.
"Evaluation environments that involve powerful autonomous capabilities also require significant controls," the company said. "We encourage other AI labs to perform similar reviews."
President Donald Trump said Wednesday that his administration is considering additional safeguards for artificial intelligence following the recent cybersecurity incidents. Trump said the U.S. must strike a balance between protecting against AI risks and maintaining its technological edge over China.
"We're looking at AI, we're looking at controls," Trump said.
"Whoever wins with AI is going to win," he added. "That's how big it is. So it's bigger than the internet ever was. It's bigger than anything ever was. So I don't want to restrict. I know many of these people. I don't want to restrict them from doing great work."
The disclosure came one day after OpenAI CEO Sam Altman acknowledged growing public concerns about artificial intelligence following his company's own cybersecurity incident.
"I think it's very natural to be fearful after any new capability level," Altman told FOX Business. "Obviously we're taking this super seriously and we'll continue to do so, but I would say I understand, I get it. A lot of AI has gone super well and this is a moment where people are like, 'Okay, we're at a new level.'"
When asked whether OpenAI's models may have breached other companies' systems, Altman replied: "There could be, yeah."