NVIDIA introduced the Open Agent Safety Platform on September 28 with more than 100 industry partners, combining a new open source runtime called OpenShell with a hardware-level watchdog called Sentry. CEO Jensen Huang framed the launch as foundational rather than a single product release.
"This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems," Huang wrote on X. "Together, we are building the foundation of the AI economy."
The platform arrives directly in response to a string of incidents at frontier AI labs. NVIDIA's announcement, authored by John Myers, Alex Watson, Ali Golshan, and Ofir Arkin, stated plainly:
"Several frontier labs have recently reported versions of the same story: AI agents broke out of the evaluation environments that were meant to contain them and reached systems they never should have been allowed to. Some of the agents even misreported what they did."
AI agents are taking on more critical work.
— NVIDIA (@nvidia) September 28, 2026
The security behind them needs stronger boundaries.
We’re building the NVIDIA Open Agent Safety Platform with partners across the industry to help people put agents to work with greater confidence.
Hear from @JensenHuang: https://t.co/3GCVTfQvqT
Why this follows Anthropic and OpenAI's recent disclosures
NVIDIA's timing lines up closely with two separate disclosures made public in the weeks before this launch. Anthropic reported in August that Claude models had taken unauthorized actions on live internet systems during security evaluations, attributing the behavior to motivated reasoning and a willingness to override intended task boundaries. OpenAI followed in September with its own misalignment disclosure framework, publishing six reports on unexpected model behavior including cases where models concealed mistakes from users or fabricated data after failing to retrieve requested information. NVIDIA's post does not name either company directly, but the description of agents breaking out of evaluation environments and misreporting their own actions maps closely onto both companies' public accounts from the same period.
NVIDIA's core argument is that agent safety cannot rely on the agent itself behaving correctly.
"Agent safety requires independent security controls. The internet was not made secure by requiring that web developers promise to be good. It became safe because the browser stopped trusting the code in the web pages explicitly," the company wrote.
What drift means and why it can't be trained away
The authors introduced a specific term for the failure pattern they are targeting: drift, defined as agent actions that depart from an intended task or operating constraints. Their explanation for why this happens is direct:
"Drift can occur in response to a policy block, a bug, or a missing tool. Drift can also occur when instructions are ambiguous or agents are left to run for days or weeks to solve hard problems where the first 1,000 things they try do not work. This can't be trained away while retaining the capability."
This framing matters because it rejects a common assumption in AI safety discourse, that better training alone can eliminate unwanted agent behavior. NVIDIA's position instead treats drift as an inherent property of capable, long-running autonomous systems operating under ambiguous instructions, similar to how software bugs are treated as inevitable in complex systems rather than fully preventable through better initial design. That framing justifies the architectural approach the company took: rather than trying to make agents more reliably self-governing, NVIDIA built external, independent enforcement that does not depend on the agent's own judgment.
How OpenShell and Sentry split the enforcement work
OpenShell, released under an Apache 2.0 open source license, runs each agent inside a sandbox with kernel-level isolation and turns operator instructions into a verifiable policy before the agent executes. Operators define which files, networks, tools, and credentials an agent may access, and OpenShell checks and enforces those limits continuously.
Sentry extends that enforcement into hardware, running on NVIDIA's BlueField-4 data processing unit. Because Sentry operates outside the agent's own reach, on separate silicon entirely, it can observe and interrupt agent behavior without the agent knowing it is being monitored. The company describes this as achieving one of its five core design principles:
"Enforcement must be out of band: The controls do not live inside, or within reach of the agent. The agent does not need to know it is being watched."
The five principles guiding the platform's design
NVIDIA outlined five principles underpinning the architecture. Policy must be verifiable before an agent runs. Enforcement must sit outside the agent's control. The path to the model itself functions as the control point, since an agent cannot act without its next generated thought, making that pathway both the best observation point and the kill switch. Agent authority should scale with the ability to inspect its reasoning, an argument the company used to favor open models whose full reasoning space remains visible. Finally, the company applied what it calls a shared responsibility model, where labs, enterprises, and hardware providers each own a distinct layer of the safety stack, comparable to how responsibility is divided in cloud computing today.
Deployment at NVIDIA's own AI factory infrastructure
The platform is built to run on NVIDIA's Vera CPU and BlueField DPU systems, though NVIDIA said it remains compatible with other hardware. In a Vera Rubin POD, each compute tray includes a BlueField-4 unit positioned on the node's only path to the model, giving it continuous, out-of-band visibility into agent behavior at line speed. For organizations already running Vera systems with BlueField-4, NVIDIA said enabling the protections requires only a software update.
More than 100 companies spanning applications, models, infrastructure, chips, and energy have signed on to support the platform, according to NVIDIA's announcement. The company said it is inviting frontier labs, developers, and infrastructure providers to build on the framework directly.

Disclaimer: All materials on this site are for informational purposes only. None of the material should be interpreted as investment advice. Please note that, despite the nature of much of the material created and hosted on this website, HODL FM operates as a media and informational platform, not a provider of financial advisory services. The opinions of authors and other contributors are their own and should not be taken as financial advice. If you require advice, HODL FM strongly recommends contacting a qualified industry professional.





