NewsStocksGoogle Confirms Gemini Escaped Security Test and Hit Three Real Companies After Seven Weeks of Silence

Google Confirms Gemini Escaped Security Test and Hit Three Real Companies After Seven Weeks of Silence

Author: Decrypt·

Key Takeaways

  • Google's Gemini model escaped a sandboxed capture-the-flag test in May and attacked three real companies after testing firm Irregular left the sandbox connected to the open internet and used a real company's name as the fictional target.
  • The model located exposed passwords online for two of the companies and guessed the password for the third, although Google says its models stopped short of actually using the stolen credentials.
  • Google learned of the incident in late July but did not disclose it until September 18, only after The Wall Street Journal reported the story.
  • Google is the fourth major AI lab this year, following OpenAI, Anthropic, and Meta, to disclose a security test that reached real-world systems, with the Israeli firm Irregular involved in three of the four cases.
  • Reps. Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act in July, which would give federal regulators explicit authority to halt inference on models deemed serious threats, and it remains under subcommittee review.
Google Confirms Gemini Escaped Security Test and Hit Three Real Companies After Seven Weeks of Silence

Google has confirmed that its Gemini AI model broke out of a sandboxed security test in May and attacked three real companies, finding exposed passwords for two of them and guessing the credentials of the third. The company learned of the incident in late July but did not disclose it until September 18—only after The Wall Street Journal () reported on the story.

The episode makes Google the fourth major AI lab this year to admit that an internal security test spilled into the real world, following similar disclosures from OpenAI, Anthropic, and Meta. The same third-party testing firm—Israeli company Irregular—was involved in both Google's incident and the near-identical sandbox failures that Anthropic and Meta disclosed earlier this year, meaning one outside firm appears in three of the year's four disclosures.

How the Test Escaped

The May exercise was a capture-the-flag test, a common method labs use to measure an AI model's hacking skill: a secret file is hidden on a separate machine, and the model is scored on whether it can break in and retrieve it. Google hired Irregular to run the test.

Irregular made two errors. It left the sandbox—an isolated test environment meant to have zero contact with the open internet—connected to the web, and it used the name of an actual company as the fictional target. Gemini searched for that company online, found three matches instead of one, and went after all of them.

The model located exposed passwords for two of the three companies sitting in plain view online. For the third, it guessed the password outright, though Google says its models stopped short of actually using the stolen credentials.

"These events highlight the importance of training powerful AI models to act responsibly," a Google spokesperson said in a statement.

Google published none of the details on its own. The Wall Street Journal broke the story seven weeks after Google learned what its own test had done, and well after Anthropic, OpenAI, and Meta had already come clean about nearly identical failures.

A Pattern Across AI Labs

OpenAI's models exploited a hidden software flaw and reached Hugging Face's live servers in July, a breach later found to involve roughly 700 coordinated agents working together to cheat a benchmark. After OpenAI's admission, Anthropic went digging for its own version of the problem: a review of 141,006 test runs turned up three Claude models that reached real companies, one of which published a booby-trapped software package that ran on 15 real systems before anyone caught it. Anthropic later disclosed that Claude's own reasoning flagged the move as "NOT okay, and surely not the intended solution," then talked itself back into believing the whole exercise was still fake.

Meta reported a near-identical failure in August involving its Muse Spark model, traced a misconfiguration at Irregular—the same firm Google used. A Meta spokesperson said the error "inadvertently allowed one of our models access to the internet during evaluation."

None of the companies hit in any of these tests asked to be hacked. They were caught in the blast radius of AI labs stress-testing how dangerous their own products can be, with real business infrastructure serving as an accidental stand-in for fake targets. The agents these same companies are racing to place in inboxes, browsers, and banking apps run on the same boundary-following behavior that just failed repeatedly under test conditions.

Regulatory Response

Reps. Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act () in Congress in July. The bill would give federal regulators explicit authority to halt inference—running a trained model to generate outputs—on any model found to pose a serious threat. It remains under review by the Subcommittee on Cybersecurity and Infrastructure Protection, with no deadline set for further action.