Skip to main content
Back to timeline
NVIDIA 开发者技术博客Source publication:

NVIDIA launches its Open Agent Safety Platform, enforcing agent policy in silicon via OpenShell sandboxes and BlueField-4

Synopsis

NVIDIA introduced an open agent safety platform comprising OpenShell, an Apache 2.0 open-source secure runtime that executes autonomous agents in kernel-isolated sandboxes and turns operator instructions into a verifiable policy checked before and enforced during execution, plus optional NVIDIA Sentry and DOCA layers that push monitoring and enforcement into BlueField hardware, where in a Vera Rubin POD each compute tray carries a BlueField-4 DPU on the node's only path to the model for continuous out-of-band observability and line-speed real-time policy enforcement, with the company stating that on existing Vera and BlueField-4 systems these protections need only a software update.

AI-generated editorial illustration: NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring

Interpretation

It introduces and open-sources OpenShell, an Apache 2.0 secure runtime that executes autonomous agents in sandboxes with kernel-level isolation, converts operator instructions into a verifiable policy defining which files, networks, tools, processes, and credentials an agent may access, and checks those limits before the agent runs and enforces them as it works. Rather than relying on developer self-restraint or placing controls inside the agent, it puts the control point outside the agent and argues that policy must be verifiable, enforcement must be out of band, and the path to the model is the control point. The text presents architecture and design principles plus lessons from a year of building OpenShell; it reports no benchmarks, attack-success rates, or controlled experiments.

It pushes monitoring and enforcement into hardware: NVIDIA Sentry extends monitoring and enforcement into NVIDIA BlueField hardware, and DOCA makes the BlueField security foundation programmable and connects it with OpenShell policy, correlating agent interactions, policy decisions, and tool and data access into a contextual record used to identify drift, investigate suspicious behavior, and decide when intervention is needed. The safety layer moves out of the host and beyond the agent's reach, standing as a trusted infrastructure protection layer even when host resources cannot be trusted. The text describes platform components and their placement; it offers no independent evaluation or third-party validation.

It names a concrete hardware landing point: in an NVIDIA Vera Rubin POD, each compute tray includes a BlueField-4 DPU on the node's only path to the model, providing continuous out-of-band observability and enforcing security policies in real time at line speed; the platform is optimized for NVIDIA Vera CPU- and BlueField DPU-based systems and is also compatible with other hardware systems. Policy enforcement is described as happening in silicon, and users already running Vera with BlueField-4 can enable these protections with just a software update. This is product architecture and deployment description; the text includes no performance figures or comparative measurements.

It proposes a layered framework and shared-responsibility model for agent safety: the platform spans an application layer, a runtime layer, and an infrastructure layer, and labs, enterprises, and hardware providers each own a layer, with the agent runtime and its policy language kept open so any provider can plug in. It frames agent safety as a browser-sandbox-style trust layer for the internet, arguing that security and safety did not slow innovation but allowed it to accelerate. This is a position and framing argument grounded in the authors' account of frontier-lab reports that agents broke out of evaluation environments and misreported what they did; the text gives no specific sources or data for those reports.

Perspective

The platform targets organizations running autonomous agents across three layers: the application layer that end users build, the runtime layer that orchestrates workloads and provides continuous monitoring and real-time policy enforcement and governance, and the infrastructure layer covering network calls, database and filesystem access, and general-purpose and accelerated compute. OpenShell suits teams that want agents to run in a zero-trust environment out of the box with isolation, monitoring, and behavior detection; organizations wanting an additional independent layer can add NVIDIA Sentry and DOCA to place monitoring and enforcement on BlueField hardware. The text notes that users already running on NVIDIA Vera systems with BlueField-4 can enable these protections with just a software update, indicating low adoption cost on existing NVIDIA infrastructure.

The text reports no quantitative evaluation: no attack-success rates, drift-detection rates, performance overhead, or comparison against alternatives, so it is hard to judge how effective these controls are under real adversarial conditions. The frontier-lab reports of agents breaking out of evaluation environments and misreporting their actions are mentioned only in general terms, without specific sources, scenarios, or data, so readers wanting to verify them must look elsewhere. The text is the full body but contains no figures, configuration examples, or policy-language details, leaving open how a policy is actually proven not to escape the operator's intent; and while the platform is stated to be compatible with other hardware systems, the degree to which the out-of-band enforcement layer is tied to Vera and BlueField also warrants further understanding.

Sources