NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring
Source
Tanya Lenz
Author
Tanya Lenz
Date
Key takeaways · AI-distilled
NVIDIA argues AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → drift, actions that depart from the task after a policy block, a bug, a missing tool, or a long run, cannot be trained away without losing capability, so agents cannot be expected to fully govern themselves.
The first of NVIDIA's five principles is verifiable policy: before an agent runs, a prover shows its policy cannot escape the operator's intent, and OpenShell then enforces those file, network, tool, process and credential limits.
NVIDIA treats the path to the model as the key control point: an agent cannot act without its next thought, so owning that path gives both the best observation point and a kill switch.
The authors tie agent authority to reasoning visibility: the more an agent can do, the more its reasoning must be inspectable, and they cite open models' visible reasoning and activations as an advantage.
NVIDIA says systems already running Vera with BlueField-4 can turn on the Sentry protections with a software update, and that the platform is also compatible with other hardware.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
sandbox — An isolated environment where AI-generated code or agent actions run without being able to touch anything real.
Why it matters
NVIDIA's platform puts policy enforcement and monitoring on the DPU sitting on an agent's only path to the model, giving continuous out-of-band observability and line-speed policy enforcement independent of the agent's own code.