Ng: sandbox agents with deterministic limits, not prompts
- Source
- Andrew Ng
- Date

The OpenAI-Hugging Face hack was enabled by weak sandboxing. It is great that Nvidia is releasing open source tools for sandboxing AI agents. OpenWorker, our open-source agent harness supporting cybersecurity workflows, is proud to support this. A sandbox gives an agent limited permissions. OpenWorker is building on Nvidia OpenShell and will support running each agent's commands inside a sandbox. Only the files relevant to the task go in. Secret API keys, your web browser login credentials, the ability to access arbitrary websites, are inaccessible to the agent by default. These restrictions are implemented in deterministic code rather than by prompting an LLM, which can make mistakes or be susceptible to prompt injections. Further, all actions are logged for monitoring and audit. I'm grateful for @JensenHuang's leadership making AI agents more secure. OpenWorker (which @rohitcprasad and I are working on) will continue to improve security for agents.
- AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
- sandbox — An isolated environment where AI-generated code or agent actions run without being able to touch anything real.
- prompt injection — An attack that hides instructions in content an AI will read — a webpage, email, or document — tricking it into following the attacker instead of the user.
Enforcing permissions in deterministic code, not prompts, keeps API keys, browser credentials and arbitrary web access away from an agent even under . It gives a concrete design principle for securing agent harnesses.
Checking sign-in…
Loading comments…





