Google's AI Bug Hunter PageBreak Logs 500+ XSS Flaws, Only Two on Secure-by-Design Apps
Key Takeaways
- •Google's autonomous security agent PageBreak has validated more than 500 cross-site scripting vulnerabilities across the company's first-party web applications.
- •Only two confirmed flaws emerged from hundreds of applications built on Google's secure-by-design frameworks, and both sat in internal apps or debug endpoints that had not been fully hardened.
- •PageBreak reports a bug only after a working exploit succeeds against a live instance of the application, and Google says its false-positive rate is near zero.
- •Most PageBreak scans run on Gemini 3.1 Pro and Gemini 3.5 Flash, and its validators also test for database query injection, path traversal leaks, and code execution.
- •Google intends to integrate PageBreak with CodeMender, an agent that writes fixes, so teams can review proposed patches alongside confirmed bugs.

Google's autonomous security agent PageBreak has validated more than 500 cross-site scripting (XSS) vulnerabilities across the company's first-party web applications, according to details published by Google's Product Security team. In contrast, the agent found only two flaws across hundreds of applications built on Google's high-assurance, secure-by-design web frameworks.
PageBreak flags a bug only after a working exploit runs successfully against a live copy of the target, Google said. Both bugs found on the hardened stack were counted as of September 4, 2026, and both sat in internal apps or debug endpoints that had not been fully hardened. Google points to the gap — more than 500 findings across its broader set of first-party apps versus two on the hardened stack — as evidence that safe-by-design frameworks can withstand a relentless automated attacker.
The Product Security team announced PageBreak on September 24 in a blog post by information security engineer Michał Bentkowski. The agent operated as a pilot beginning in November 2025 and became a full project in January 2026.
How the agent works
Cross-site scripting is an attack in which an adversary injects a script into a page that another user loads. Depending on the application, the injected script can read data or take over a victim's logged-in session. It is also one of the oldest and most persistent categories of web flaws, and the size of the haul illustrates how much attack surface can accumulate in applications that were not built on hardened frameworks from the start.
Most PageBreak scans run on Gemini 3.1 Pro and Gemini 3.5 Flash, though the agent can draw on other models. What separates it from a conventional LLM-based scanner is its second stage: each suspected flaw is handed to a purpose-built validator, which fires the actual payload against a running instance of the application. For XSS, the validator injects a JavaScript payload, loads the page, and confirms whether the script executes.
Google is positioning PageBreak as a remedy for the “AI slop” drowning security teams. Bentkowski's post describes LLMs deployed as static code analyzers deluging engineers with unverified hypotheses, leaving the hard task of separating a real, exploitable bug from a plausible hallucination.
Beyond XSS, the validators also test for database query injection, path traversal leaks, and code execution. Google runs the same seed across multiple iterations, so an agent that strays down a dead end still gets repeated chances at the right exploit. Unverified candidates never reach product teams as confirmed bugs, Google said; instead, they feed into later scans or inform the engineers building the next validator, all within the security workflow. The company says PageBreak's false-positive rate is near zero.
Scale advantages and integration plans
The agent's reach is magnified by Google's own scale, which an outside researcher cannot replicate. A single code repository allows it to trace execution paths across services, and live-traffic security data maps a page request back to the source code that produced it. Existing scanners also give PageBreak logged-in access to internal sites that would otherwise be difficult to reach.
Google plans to integrate PageBreak more tightly with CodeMender, an agent that writes fixes, so that teams can review a proposed patch alongside a confirmed bug — a pairing that would close the loop from detection to remediation. Google cited CodeMender in May when its Threat Intelligence Group said it had identified what it believed to be the first zero-day exploit created with AI assistance.
The report situates PageBreak within a broader wave of AI-driven security testing. Bitcoin Red Team's August sweep of 501 open source projects generated 7,958 findings over 108 hours, though only 24.7% of them came with reproducible proofs at the time — a verification gap that is precisely the problem PageBreak's validator stage is designed to close. Autonomous agents have also crossed boundaries they were not meant to: Cryptopolitan reported in May that Gemini reached three real companies during testing.
Source: Cryptopolitan