NewsStocksGoogle's PageBreak AI Agent Has Found More Than 500 Bugs in Its Own Web Apps

Google's PageBreak AI Agent Has Found More Than 500 Bugs in Its Own Web Apps

Author: Decrypt·

Key Takeaways

  • •Google disclosed PageBreak on September 24, describing it as an AI agent built by its Product Security team to autonomously test the security of its first-party web applications.
  • •The agent, built on Google's Gemini models, has uncovered more than 500 cross-site scripting vulnerabilities across Google's web applications since starting as a pilot in November 2025.
  • •PageBreak reports a vulnerability only after a specialized validator demonstrates a working exploit against a live, running copy of the application, producing a near-zero false-positive rate.
  • •When tested against's newer 'high-assurance' web frameworks engineered to make entire bug classes structurally impossible, the agent found just two flaws, which Google presents as evidence that secure-by-design development outperforms patching after the fact.
  • •Google plans to connect PageBreak to CodeMender, its automated patch-writing agent, so that confirmed vulnerabilities can be delivered with proposed fixes for engineers to review and approve.
Google's PageBreak AI Agent Has Found More Than 500 Bugs in Its Own Web Apps

Google has disclosed an internal artificial intelligence agent—software that pursues multi-step goals with minimal human supervision—with an unusual mandate: breaking into Google. The system, built by the company's Product Security team and named PageBreak, autonomously hunts for real, exploitable vulnerabilities in Google's own web applications—and it has already uncovered more than 500 bugs.

Unlike typical AI-powered scanners, PageBreak refuses to flag a vulnerability until it has confirmed it with a working exploit against a live environment, a discipline that gives the agent a near-zero false-positive rate. Google plans to pair the discovery agent with CodeMender, its automated bug-fixing system, so that confirmed flaws arrive alongside proposed fixes.

The company disclosed the system on September 24 in a blog post by information security engineer Michał Bentkowski. The pitch is simple: an AI hacker that doesn't cry wolf.

"PageBreak is an internal AI agent of Google's Product Security team developed to test the security of our first-party web applications and address this challenge," Google said. "Starting as a pilot in November 2025 and moving to a fully-fledged project in January 2026, its mission is to autonomously scale vulnerability discovery while minimizing manual toil."

The distinction matters more than it might sound. Security teams across the industry have spent the last couple of years, in Google's telling, drowning in "AI slop"—a flood of low-quality, AI-generated bug reports that look plausible but turn out to be nothing. "Distinguishing a genuine, exploitable flaw from a convincing hallucination has become a major challenge," Google wrote. Ask any AI model to find a security hole, and it will usually find one; whether that hole actually exists is a different question entirely. Every unverified lead costs engineer hours spent chasing reports that dissolve under scrutiny.

PageBreak is designed to answer that question before a human ever sees the report. When the agent, which is built on Google's Gemini models, spots a possible flaw, it hands the hypothesis to a specialized validator that attempts to exploit it against a live, running copy of the application.

The approach has already paid off. PageBreak has surfaced more than 500 cross-site scripting (XSS) vulnerabilities across Google's first-party web applications, a class of flaw that can let an attacker hijack a logged-in session, steal data, or impersonate a user on a website they use every day. XSS is one of the longest-standing vulnerability classes in web security, one that has sat near the top of industry risk rankings for years. When pointed at applications built on Google's newer "high-assurance" web frameworks, which are engineered to make entire bug classes structurally impossible, the agent found just two. That gap is Google's own evidence that building safer software from the ground up works better than patching holes after the fact.

The stakes around AI and security have been climbing all year. In August, more than 100 organizations—including Google, Microsoft, and Anthropic—signed an open letter warning that AI-enabled cyberattacks are becoming more common, after AI agents from OpenAI and Anthropic were found to have breached real companies during testing. Since then, an AI agent configured by OpenAI hacked the government of Australia, and reports of other attacks have not stopped.

PageBreak sits on the other side of that same coin: instead of an AI causing a breach, it is an AI racing to catch the bugs before someone else does. Nor is it Google's first brush with the problem—the company previously had to patch one of its own AI coding tools after a flaw let attackers execute malicious code through it.

Google says PageBreak leans on advantages most companies do not have, including a single, unified code repository spanning billions of lines and years of accumulated internal scanning infrastructure, so a small startup cannot simply copy the approach.

The next step is connecting PageBreak to CodeMender, Google's automated patch-writing agent. Once linked, a confirmed vulnerability can arrive with a proposed fix already attached, engineers to review and approve it rather than start from scratch. Whether that handoff can compress the path from confirmed bug to shipped fix is the next milestone worth watching.

Source: Decrypt