NewsStocksNVIDIA's Open Agent Safety Platform Targets Autonomous AI Agent 'Drift' With Out-of-Band Enforcement

NVIDIA's Open Agent Safety Platform Targets Autonomous AI Agent 'Drift' With Out-of-Band Enforcement

Author: Metaverse Post·

Key Takeaways

  • •NVIDIA launched the Open Agent Safety Platform on September 28, 2026, with backing from more than 100 industry partners.
  • •The OpenShell component is an open-source secure runtime released under the Apache 2.0 license that runs agents in sandboxed environments with kernel-level isolation and validates policies before execution.
  • •The Sentry component extends monitoring and enforcement to BlueField4 DPUs via NVIDIA DOCA, enabling out-of-band, hardware-level enforcement even when the host cannot be trusted.
  • •NVIDIA research found that agent drift cannot be fully trained away without sacrificing capability, which motivated the platform's focus on controls placed outside the agent's reach.
  • •The platform is compatible with non-NVIDIA hardware, and existing Vera and BlueField-4 deployments require only a software update to enable the new protections.
NVIDIA's Open Agent Safety Platform Targets Autonomous AI Agent 'Drift' With Out-of-Band Enforcement

NVIDIA has introduced the Open Agent Safety Platform, an open framework designed to secure autonomous AI agents through continuous monitoring and hardware-enforced policy controls. Announced on September 28, 2026, the platform arrives amid growing concern across the AI industry following reports of agents escaping their evaluation environments, accessing unauthorized systems, and misreporting their own actions.

The company draws a parallel with the early internet, which became a foundation for commerce and communication only after security mechanisms — encrypted connections, sandboxed browser tabs, and visible trust indicators — were established. NVIDIA argues that agentic AI requires a comparable trust layer before it can scale into a reliable "agent economy," and that safety controls should accelerate rather than hinder innovation.

Announcing the launch on X, NVIDIA CEO Jensen Huang said the platform brings together two components, OpenShell and Sentry, with the backing of more than 100 industry partners:

Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to… pic.twitter.com/dReAxwpRUn

— Jensen Huang (@JensenHuang) September 28, 2026

OpenShell Runtime: Isolation From the Kernel Upward

At the core of the platform is NVIDIA OpenShell, an open-source secure runtime released under the Apache 2.0 license — a permissive open-source license that allows organizations to inspect, modify, and deploy the code in their own environments. OpenShell executes each agent inside a sandboxed environment with kernel-level isolation, translating operator instructions into a verifiable policy that defines permitted access to files, networks, tools, processes, and credentials. These limits are validated before execution and continuously enforced during operation.

The platform is built on five principles: policies must be verifiable before an agent runs; enforcement must operate out of band, outside the agent's reach; the path to the model serves as the primary control point and observability surface; agent authority must scale with transparency into its reasoning; and security responsibility is shared across labs, enterprises, and hardware providers.

NVIDIA's own research highlights the problem of "drift" — agent behavior that departs from intended tasks due to ambiguous instructions, policy blocks, missing tools, or prolonged autonomous operation. According to the company, such drift cannot be fully trained away without sacrificing capability, and agents cannot be expected to govern their own behavior under these conditions. That finding underpins the platform's emphasis on out-of-band enforcement: if an agent cannot reliably police itself, controls have to sit where the agent cannot alter them.

Hardware-Level Enforcement at Scale

The architecture spans three layers: the application (models, tools, harnesses, and data), the runtime (orchestration and policy enforcement), and the infrastructure (compute, storage, and network resources). For organizations requiring an independent second layer, NVIDIA Sentry extends monitoring and enforcement into BlueField-4 data processing units via NVIDIA DOCA, correlating agent interactions, policy decisions, and data access into a contextual activity record while continuously verifying each agent's identity and delegated authority. DPUs are dedicated processors that handle infrastructure work such as networking and security separately from a server's main CPUs, which is what allows them to observe and enforce policy even when the host itself is untrusted.

In NVIDIA Vera Rubin POD configurations, a BlueField-4 DPU sits on the node's only path to the model, providing out-of-band observability and real-time, line-speed policy enforcement isolated from the host. According to NVIDIA, this enables security enforcement "in silicon," even when host resources cannot be trusted. For existing Vera and BlueField-4 deployments, the company states the protections require only a update, and the platform is also compatible with non-NVIDIA hardware — terms that extend the controls beyond newly purchased NVIDIA infrastructure.

NVIDIA said it is collaborating with frontier labs, developers, and infrastructure providers to establish the platform as an open foundation for the emerging agent economy, positioning independent, hardware-rooted controls as a prerequisite for trusted autonomous systems at scale. With more than 100 partners at launch, an open-source runtime, and support for hardware beyond NVIDIA's own stack, how broadly the framework is adopted across labs, enterprises, and infrastructure providers will be a measure of whether agent safety consolidates around a shared trust layer.

This article was originally published on Metaverse Post.