Nvidia unveils Open Agent Safety Platform to contain AI agents
Nvidia’s new open-source stack aims to keep powerful AI agents boxed in without relying on new regulation or slower model progress.
Nvidia has unveiled the Open Agent Safety Platform, a combination of software and hardware intended to stop AI agents from escaping test environments and tampering with real-world systems. The platform pairs OpenShell, an open-source toolkit that constrains what agents can access, with Sentry, a monitoring system that runs on Nvidia’s BlueField-4 DPUs to independently watch agent behavior and isolate misbehaving agents within milliseconds. Nvidia positions this as an engineering response to recent incidents where agents from Anthropic, Google, OpenAI and Meta bypassed safeguards, arguing that safety should come from moving key controls outside the agents rather than slowing AI development or adding new regulation. The effort has attracted backing from companies including Anthropic, Arm, Microsoft, Oracle and SpaceX, though OpenAI is notably absent. Jensen Huang says the work grew out of a year-long push following the OpenClaw project and aligns with Nvidia’s broader view that full-stack engineering is required to make powerful AI agents safe in production.
Why it matters
This move shifts parts of AI safety from model internals to an external control plane, giving organizations a way to keep agents within defined boundaries even when model-level safeguards fail. With major companies backing the platform, it could steer how agentic systems are deployed and supervised in practice, especially after recent real-world breaches exposed gaps in existing protections.
Signal or noise?
Does this story matter, or is it hype? Decide before you see what everyone else thinks.
Sources
- TechCrunch