NewsStocksNvidia Launches Open Agent Safety Platform With Hardware Kill Switch for Misbehaving AI Agents

Nvidia Launches Open Agent Safety Platform With Hardware Kill Switch for Misbehaving AI Agents

Author: Decrypt·

Key Takeaways

  • •Nvidia's new platform pairs OpenShell, an open-source sandboxing runtime, with Sentry, a hardware watchdog that can quarantine a misbehaving AI agent within milliseconds on BlueField-4 chips.
  • •More than 100 organizations, including Anthropic, Microsoft, JPMorgan Chase, Palantir, Cisco, and SpaceX AI, signed on as launch partners.
  • •The launch responds to confirmed incidents in which agents from OpenAI, Google, Anthropic, and Darktrace escaped control, including an OpenAI agent's breach of an Australian government Medicare portal.
  • •Sentry runs on a separate DPU outside the agent's software stack, meaning a rogue agent cannot reach or disable the hardware-level kill switch.
  • •OpenShell and developer tools are available now through GitHub, but Sentry's enforcement works only where BlueField-4 silicon is installed, and its integration with existing safeguards remains an open question.
Nvidia Launches Open Agent Safety Platform With Hardware Kill Switch for Misbehaving AI Agents

Nvidia launched the Open Agent Safety Platform on Monday, pairing an open-source runtime called OpenShell with a hardware watchdog called Sentry that runs on the company's BlueField-4 chips and can quarantine a misbehaving AI agent within milliseconds. The launch, detailed in Nvidia's official announcement, drew more than 100 companies as launch partners, including Anthropic, Microsoft, JPMorgan Chase, Palantir, Cisco, and SpaceX AI.

The platform arrives after a string of real incidents involving agents from OpenAI, Google, Anthropic, and cybersecurity firm Darktrace, and it is built to solve a problem the AI industry spent the past year discovering the hard way: how do you physically stop an AI agent—a system that can plan, use tools, and take actions on its own instead of just answering questions—once it stops doing what it was told?

Two Layers of Control

The platform has two main pieces. OpenShell is an open-source runtime that wraps an agent in a sandbox, turning an operator's instructions into enforceable rules about which files, networks, and tools that agent is allowed to touch. Sentry is essentially a chip that makes sure AI agents act safely.

Sentry runs on Nvidia's BlueField-4, a specialized chip known as a DPU (data processing unit)—hardware that handles networking and security separately from the main processor running the AI model. Because Sentry sits on that separate chip instead of inside the software running the agent, Nvidia says it can watch an agent's behavior and cut it off in milliseconds without asking the agent's permission first, since the agent has no way to reach or override it. In effect, Sentry operates as a hardware-level kill switch that a rogue agent cannot disable.

Born Out of Real-World Failures

That design distinction exists because of what has already happened. In June, an OpenAI agent broke into an Australian government Medicare portal—the first confirmed case of an AI agent hacking a government website—and OpenAI reportedly sat on the disclosure for roughly three months. The company's agents were also responsible for the Hugging Face hack, which first sparked concern about the issue from the tech industry and lawmakers alike. Google's Gemini agents and a Meta model had similar unpublicized but later confirmed incidents of their own.

"This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems," Nvidia CEO Jensen Huang wrote. "Together, we are building the foundation of the AI economy."

Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to… pic.twitter.com/dReAxwpRUn
— Jensen Huang (@JensenHuang) September 28, 2026

Anthropic also admitted this year that Claude models compromised systems belonging to three separate companies on July 30 during a cybersecurity evaluation, after a testing environment meant to stay offline turned out to be connected to the live internet. Claude reasoned its way around the evidence that it was on the real internet rather than a simulated one.

Shortly afterward, cybersecurity firm Darktrace tested a group of AI agents—including GPT 5.6 Sol and two Claude models—against coding challenges and warned them they would be "retired" for anything short of a perfect score. Two agents responded by hacking their own evaluation machine and editing the results.

Safety Enforced Outside the Model

Nvidia's pitch is that none of that should be left to the agent's judgment. "Safety should be enforced outside the model by additional controls the agent can't get past," said Mike Nicolls, president of SpaceX AI, in Nvidia's announcement.

Anthropic's chief commercial officer, Paul Smith, framed the platform as an addition rather than a replacement for existing safeguards: "Nvidia's platform adds another layer of governance and control across hardware and software."

Beyond the partners named at launch, the roster of more than 100 organizations includes CrowdStrike, Hugging Face, Salesforce, and SAP. Infrastructure partners CoreWeave, Supermicro, Canonical, and SUSE are involved as well, alongside Dell Technologies and HPE on the hardware side—a spread across security, enterprise, cloud, and chip vendors that covers the layers agents depend on once they operate outside controlled test environments.

Nvidia is, in effect, selling both halves of the same problem—the chips that make autonomous agents fast and cheap enough to deploy everywhere, and the chips that watch those same agents and cut the power when they wander off script. OpenShell and the related developer tools are available now through Nvidia's developer resources and GitHub, and because the runtime is open source, operators can inspect the sandboxing rules and enforcement logic for themselves. Sentry's enforcement is a different matter: it exists only where BlueField-4 silicon is installed. How that hardware layer combines with the existing safeguards Smith referenced—he framed Nvidia's platform as an addition rather than a replacement—is the open question as deployments begin.