↓ Skip to main content

NVIDIA Open Agent Safety Platform: Hardware-Enforced AI Security

NVIDIA Open Agent Safety Platform: Hardware-Enforced AI Security

As AI agents become capable of executing code, accessing files, calling APIs, and operating enterprise systems autonomously, security boundaries can no longer depend entirely on the model or agent itself.

NVIDIA has introduced the Open Agent Safety Platform, an open software platform and reference system design that combines runtime controls with hardware-isolated enforcement. The architecture brings together NVIDIA OpenShell and NVIDIA Sentry, with Sentry running on NVIDIA BlueField-4 DPUs as an independent, out-of-band security layer.

The key architectural idea is simple: the agent should not be responsible for enforcing the rules that constrain the agent.

🛡️ OpenShell: The Runtime Security Boundary
#

NVIDIA OpenShell provides the software layer of the platform. It runs agents inside isolated sandboxes and enforces security policies outside the agent process.

Rather than relying solely on system prompts or model alignment, OpenShell establishes explicit boundaries around what an agent can access and do.

Its security model includes:

  • Sandboxed execution — Each agent runs in an isolated environment with controlled filesystem and network access.
  • Least-privilege execution — Agents run without unnecessary host privileges.
  • Kernel-level controls — OpenShell monitors and filters system calls and restricts unsafe operations.
  • Credential isolation — Credentials are brokered through the runtime rather than exposed directly to the agent.
  • Network policy enforcement — Network requests are evaluated against configured policies before access is granted.
  • Auditability — Allow and deny decisions are recorded for security investigation and compliance.
  • Policy verification — OpenShell includes a policy prover that uses formal verification to check whether policies remain within defined access boundaries.

OpenShell is open source under the Apache 2.0 license and is designed to work across different agent frameworks, models, and infrastructure environments. NVIDIA also states that it can be extended to third-party compute platforms, including Arm and Intel systems.

Why Put Security Outside the Agent?
#

An autonomous agent can generate code, invoke tools, modify files, and make network requests. If the same software stack is responsible for both executing the task and deciding whether its own actions are safe, a compromised or misbehaving agent may attempt to circumvent those controls.

OpenShell instead places policy enforcement in the runtime boundary.

The distinction is important:

Traditional approach

Model
  │
  ├── System prompt
  ├── Alignment
  └── Agent logic
        │
        ▼
     Actions


OpenShell approach

Model
  │
  ▼
Agent
  │
  ▼
Sandbox ──► Supervisor ──► Policy
  │              │
  │              └── Credentials / Network / Files
  ▼
Controlled execution

The agent can request an operation, but the runtime decides whether that operation falls within the permitted policy.

🔒 NVIDIA Sentry: An Independent Hardware Enforcement Layer
#

OpenShell establishes the software boundary, while NVIDIA Sentry adds an independent hardware enforcement layer.

Sentry is an out-of-band watchdog implemented using NVIDIA BlueField-4 DPUs. Because it operates outside the agent and host software path, the agent cannot simply modify the monitoring component or disable its enforcement logic. ([NVIDIA Investor Relations][1])

The architecture can therefore be viewed as two separate security domains:

┌─────────────────────────────────────────────┐
│              Agent Workload                 │
│                                             │
│  Model → Agent → Tools → Files → APIs      │
└──────────────────┬──────────────────────────┘
                   │
            OpenShell Runtime
                   │
        ┌──────────▼──────────┐
        │ Sandbox + Supervisor│
        │ Policy Enforcement  │
        └──────────┬──────────┘
                   │
             Host / Compute
                   │
        ┌──────────▼──────────┐
        │   BlueField-4 DPU   │
        │      NVIDIA Sentry  │
        │                      │
        │ Out-of-band Security │
        │     Enforcement      │
        └──────────────────────┘

NVIDIA describes Sentry as a hardware-isolated security layer capable of continuously monitoring agent activity and quarantining agents that move outside their defined boundaries, with intervention occurring in milliseconds. ([NVIDIA Newsroom][2])

Sentry uses NVIDIA DOCA to provide programmable inspection, identity verification, telemetry, and zero-trust policy enforcement for data, tools, APIs, and services. ([NVIDIA Investor Relations][1])

🧱 From Software Alignment to Full-Stack Enforcement
#

The platform represents a broader change in how autonomous AI systems can be secured.

Traditional AI safety mechanisms often operate at the model or application layer:

Prompt
  ↓
Model Alignment
  ↓
Agent Logic
  ↓
Tool Execution

NVIDIA’s architecture adds enforcement boundaries underneath the application:

┌──────────────────────────────┐
│ Application / Agent          │
├──────────────────────────────┤
│ OpenShell Runtime            │
│ Sandbox + Policy Enforcement │
├──────────────────────────────┤
│ Host / CPU                   │
├──────────────────────────────┤
│ BlueField-4 / Sentry         │
│ Hardware Security Boundary   │
└──────────────────────────────┘

This does not eliminate model-level safety techniques. Instead, it creates additional layers that remain effective when an agent behaves unexpectedly.

NVIDIA describes the architecture around several principles, including verifiable policy, out-of-band enforcement, control over the path to the model, and independent monitoring of agent behavior. ([NVIDIA Developer][3])

🤖 Extending Agent Security to Physical Systems
#

The implications become more significant when autonomous agents move beyond software.

A coding agent that accesses the wrong file may cause a data breach. An enterprise agent that invokes an unauthorized API could trigger an incorrect transaction. A robotics agent that sends an unintended command can potentially affect the physical environment.

This creates a progression:

AI Agent
   │
   ├── Files
   ├── Network
   ├── APIs
   ├── Databases
   └── Physical Devices
              │
              ▼
       Real-world effects

For robotics and other cyber-physical systems, security therefore needs to extend beyond model behavior into the infrastructure executing the agent.

NVIDIA positions the Open Agent Safety Platform as a full-stack architecture spanning software, compute infrastructure, and robotics systems. ([NVIDIA Investor Relations][1])

🌐 Ecosystem and Industry Adoption
#

NVIDIA says more than 100 organizations are working with technologies from the Open Agent Safety Platform ecosystem. The announced participants include companies such as Anthropic, Microsoft, Salesforce, SAP, Palantir, JPMorganChase, CrowdStrike, Red Hat, Figure, Cisco, Dell Technologies, HPE, Hugging Face, ServiceNow, and others. ([NVIDIA Investor Relations][1])

The ecosystem is also extending beyond NVIDIA’s own software stack.

For example, NVIDIA and Anthropic have worked on additional security boundaries for managed agents, while Salesforce has integrated OpenShell with Slack for agent activity visibility and permission workflows. SAP is integrating OpenShell with its Joule Studio runtime and contributing engineering work toward interoperability standards. ([NVIDIA Investor Relations][1])

🤝 Open Secure AI Alliance and Linux Foundation
#

The hardware and runtime architecture is accompanied by a broader industry effort around open AI security.

The Open Secure AI Alliance has joined the Linux Foundation, providing a neutral governance structure for organizations working on open AI security tools, research, and shared defenses. ([Linux Foundation][4])

By August 2026, NVIDIA said the alliance had grown to more than 120 organizations. Its work includes efforts such as the Shared AI Findings Exchange (SAFE) guidelines, intended to turn agentic AI security incidents into reusable defensive knowledge. ([NVIDIA Blog][5])

This creates an ecosystem with multiple layers:

┌─────────────────────────────────────────┐
│ AI Applications / Agent Frameworks      │
├─────────────────────────────────────────┤
│ OpenShell Runtime                       │
├─────────────────────────────────────────┤
│ Hardware Enforcement / BlueField-4      │
├─────────────────────────────────────────┤
│ Open Security Standards & Research      │
└─────────────────────────────────────────┘

The objective is not simply to make one agent secure, but to establish reusable security primitives that can be applied across different models, agent frameworks, enterprises, and infrastructure.

⚙️ OpenShell Without BlueField-4
#

An important distinction is that OpenShell does not require BlueField-4.

NVIDIA states that OpenShell can run on supported local, cloud, on-premises, and Kubernetes infrastructure without BlueField-4. BlueField-4 and Sentry provide an additional, independent hardware security layer for environments that require stronger isolation. ([NVIDIA][6])

This makes the architecture more modular:

OpenShell only
    │
    └── Runtime sandbox + policy enforcement

OpenShell + BlueField-4
    │
    ├── Runtime sandbox
    ├── Policy enforcement
    └── Independent hardware monitoring

The distinction is useful for deployment planning because organizations can adopt runtime-level controls without immediately requiring the complete NVIDIA hardware stack.

🧠 Conclusion
#

NVIDIA’s Open Agent Safety Platform introduces a security architecture in which AI agents do not have to be trusted to enforce their own boundaries.

OpenShell provides the software enforcement layer through sandboxing, policy controls, credential isolation, network restrictions, and auditing. Sentry extends those controls into an independent hardware trust domain using BlueField-4, providing out-of-band monitoring and enforcement that remains outside the agent’s control. ([NVIDIA][7])

The more important architectural idea is broader than any individual NVIDIA product:

AI agent security increasingly needs to be enforced by the infrastructure surrounding the agent, not only by the model running inside it.

As agents gain access to enterprise data, software tools, APIs, and eventually physical machines, this separation between what an agent wants to do and what the infrastructure permits it to do may become a fundamental design principle for secure agentic systems.

Related