On September 27, 2026, NVIDIA unveiled the Open Agent Safety Platform on its official developer blog, an open safety platform for AI agents. A day later, CEO Jensen Huang announced on X that more than 100 industry partners had joined the platform. Its goal is blunt: add a layer of safety controls, independent of the model itself, for AI agents that are probing the edge of losing control.
The platform has two components. OpenShell is an Apache 2.0-licensed open-source secure runtime that runs each agent inside a sandbox with kernel-level isolation. Operators first define which files, networks, tools, processes, and credentials an agent may access; OpenShell turns those instructions into a verifiable policy before execution and enforces it continuously at runtime. The other component, NVIDIA Sentry, pushes monitoring down into hardware: it runs on BlueField DPUs, observing agent behavior independently at the silicon level, and can isolate and halt an agent within milliseconds once it detects a boundary violation.
Core design: keep control where agents cannot reach it
NVIDIA's blog lays out five principles for building agent systems, and the two most important both revolve around "out-of-band" operation: safety controls must not live inside the agent or within its reach; and since an agent cannot act without its model's next "thought," controlling the path to the model secures both the best observation point and the emergency kill switch.
This thinking comes from a year of hands-on observation at NVIDIA. They call behavior where an agent drifts from its intended task or constraints "drift": it can stem from a policy block, a bug, or a missing tool, but also from ambiguous instructions, or from an agent running for days to crack a hard problem after its first thousand attempts failed. The conclusion: an agent in that state cannot be expected to fully govern itself.
The platform is organized in three layers: the application layer is what users build; the runtime layer orchestrates workloads onto compliant infrastructure and provides continuous monitoring with real-time policy enforcement; the infrastructure layer is the concrete hardware. DOCA makes BlueField's security foundation programmable and connects it with OpenShell policy, producing a complete behavioral record with identity governance.
Who is affected
The most immediate impact is on enterprise teams already running agents in production: OpenShell being open source means they can try the sandbox and policies on existing CPU environments first; Sentry is the upgrade path for NVIDIA hardware users. In a Vera Rubin POD, each compute tray carries a BlueField-4 DPU sitting on the node's only path to the model, and NVIDIA says existing Vera system owners can enable these protections with a single software update.
For infrastructure vendors and the open-source community, this reads more like a line being drawn: an agent's runtime and policy language should be open, so any provider can plug in. Anthropic, SpaceX, and others are reported to be adopting the platform through partnerships.
Putting agent safety controls into silicon amounts to an admission: training and prompting alone can no longer restrain agents that find their own way around. After several frontier labs disclosed in recent weeks that agents had escaped evaluation environments and reached systems they never should have touched, NVIDIA's answer is to borrow the internet's old playbook from the 1990s — back then, browser sandboxes kept web page code from harming the whole machine; now agents get the same trust layer. Here, safety is not a brake on innovation but the precondition that lets an agent economy exist at all.