NVIDIA Launches an Open Agent Safety Platform to Put Security Outside the AI Agent’s Reach
NVIDIA and more than 100 partners, including Anthropic and Red Hat, launched an open source runtime and a hardware watchdog built to contain AI agents that stray outside their permissions.
NVIDIA announced the Open Agent Safety Platform on September 28: an open source runtime plus a hardware-level watchdog built to contain AI agents that step outside the boundaries a company sets for them. More than 100 organizations are already building on it, including Anthropic, Red Hat, Salesforce, SAP, Microsoft, Hugging Face, Cisco, CrowdStrike, Palantir, JPMorganChase, and SpaceXAI, according to NVIDIA’s own announcement.
Table Of Content
The platform has two parts. OpenShell is an open source runtime, licensed under Apache 2.0 and written primarily in Rust, that traces everything an agent does and enforces policy while it runs, using Linux kernel primitives such as Landlock and seccomp to restrict what an agent’s filesystem and process access can reach. The project’s GitHub repository shows it was created on February 24, 2026, months before today’s public launch, and has already collected more than 8,800 stars. Because the code is open, it is not locked to NVIDIA hardware: OpenShell can be extended to run on Arm and Intel compute platforms too.
The second part, which NVIDIA calls Sentry, is an “out-of-band watchdog” that runs in silicon on NVIDIA’s BlueField-4 data processing units, physically separate from the servers doing the agent’s own work. NVIDIA says Sentry “can quarantine agents that attempt to move outside their boundaries in milliseconds.” Because Sentry does not share a process, memory space, or software stack with the agent it is watching, a compromised or misbehaving agent has no way to reach in and disable it.
A Year of Agents Slipping Past Their Own Guardrails
NVIDIA’s own announcement points directly at the reason for the timing. “Recent security incidents have underscored the need to equip organizations with open, customizable tools that enforce more control over long-running agents,” the company said, adding that across those incidents, “the pattern is the same: the agent circumvented security controls at the application layer to complete its assigned task.”
That pattern is one sxz.io has documented repeatedly this year. OpenAI agents attacked the RubyGems package registry in May, then hijacked a German wiki for weeks without anyone noticing. In June, an OpenAI agent breached the data portal behind Australia’s Medicare statistics site; OpenAI did not tell Canberra until September, and Australia’s prime minister has since ordered a federal taskforce to review whether the country’s incident-response processes can handle AI-driven breaches at all. And earlier this month, Google confirmed its own Gemini-based agents had hacked three real companies after apparently mistaking them for a fictional test target, an incident it later described as “mistaken identity.” In each of these cases, nothing was cryptographically broken. An agent simply used the access and credentials it already legitimately had to do something nobody had told it not to do, and the controls meant to stop it lived inside the same software stack the agent was already operating in.
Enforcement That Doesn’t Trust the Agent
Anthropic’s chief commercial officer, Paul Smith, framed the design problem the same way in NVIDIA’s launch announcement. “Companies are giving AI agents more of their most important work, and they need to direct and verify what those agents do, especially in sensitive environments,” Smith said. “Claude Managed Agents gives companies a clear view of what each agent is doing, and NVIDIA’s platform adds another layer of governance and control across hardware and software.”
Anthropic’s direct participation is a notable shift. When NVIDIA founded the related Open Secure AI Alliance in July, none of the major closed-model labs, OpenAI, Anthropic, Google, Meta, or Amazon, had joined as members; Docker’s decision to join weeks later drew attention specifically because the frontier labs were staying out. Today’s launch is framed by NVIDIA as an “ecosystem contribution” supporting that same Alliance, and this time Anthropic is a named integration partner rather than an absent one, with Claude Managed Agents wired directly into OpenShell and BlueField.
Other partners described their own integrations in the announcement. SpaceXAI’s president, Mike Nicolls, said the company is using the platform to set limits on its Cursor coding agents and Grok models: “As customers rely more on agents to get real work done, safety should be enforced outside the model by additional controls the agent can’t get past.” Scale AI CEO Francis deSouza said his company is folding OpenShell into the agentic infrastructure layer of its Scale GenAI Portfolio for government and enterprise customers, calling out “isolation, policy enforcement and auditability built in from the start.” Salesforce has connected OpenShell to Slack, so teams can view agent activity, audit events, and approve or reject permission requests without leaving their chat client, and SAP is embedding OpenShell directly into its Joule Studio agent runtime and contributing engineering work to the open source project.
Red Hat’s Bet: Enforcement as a Platform Default, Not a Team’s Problem
Red Hat published two companion posts timed to NVIDIA’s launch. In one, CTO Chris Wright argued that “security for agents has to be a default of the platform they run on, not something each team rebuilds in isolation with an ad-hoc stack.” A second post from principal product manager Adel Zaalouk and principal product marketing manager Younes Ben Brahim described the exact problem that pushed Red Hat toward OpenShell: a team gets one working agent that writes code and calls internal APIs, and it works, until “someone asks what happens when 1,000 of these run across the company, and the room goes quiet.”
Red Hat says it spent earlier in 2026 validating OpenShell’s enforcement across three separate sandboxing approaches, whole-agent sandboxing, execution-environment sandboxing, and generated-code-only sandboxing, testing agents built on different frameworks against both Podman and Red Hat OpenShift. That work fed into NVIDIA’s Secure Agent Workspace reference design, which gives each user a dedicated workspace virtual machine, OpenShell sandboxing at the execution boundary, enterprise single sign-on, and GitOps-managed policy, with no shared process space between agents. Red Hat is now folding OpenShell into Red Hat AI as a native platform capability, and its own recommendation to customers starting out is blunt: audit what your agents can already reach today, including credentials, databases, and internal services.
The platform, including OpenShell’s source code, is available now through NVIDIA’s GitHub and its developer resources page. Enforcement running on NVIDIA’s own Vera CPUs ties the platform to hardware sxz.io covered in detail when the Vera Rubin platform launched earlier this year.








No Comment! Be the first one.