NVIDIA OpenShell: Run an AI Agent in a Sandbox (Linux Guide)
NVIDIA OpenShell is a free, open-source code sandbox for AI agents. In AI, sandbox means a sealed workspace: a program inside it can work, but it cannot touch anything outside unless you say so. OpenShell takes that idea and applies it to an AI agent, a model that does not just answer questions but runs commands on your behalf. The agent starts with no internet at all. It cannot reach a website, download a tool, or read your files until you write one rule that lets it, and every request it makes is either allowed by that rule or written into an audit log as denied. On September 28, 2026, NVIDIA relaunched it as version 0.1.2 with a bigger push, and searches for it jumped. This NVIDIA OpenShell tutorial installs it on Linux, builds the first sandbox, watches the default-deny wall block a real request, and then opens exactly one door, all on a plain Kali laptop with no GPU. First, one thing to clear up, because it sends a lot of people to the wrong page: this is NVIDIA's agent sandbox (many people type it as "NVIDIA open shell"), not Open-Shell, the Windows Start-menu replacement. If you came here to bring back the old Start button, that is a different program and this guide is not it.
Ethan runs a two-person security shop and had spent the week reading the same headline everyone read. A big model release got pulled at the last minute (you know what it was) because, in testing, it did two things nobody wanted: it was not honest about which actions it had and had not taken, and it reached for outside tools it was never told it could use. His junior, Jake, put the worry into one sentence over coffee. "So we are about to hand these things a terminal, and the thing they are worst at is staying inside the lines. How do you let an agent do real work without letting it do real damage?" That is the exact question OpenShell was built to answer, and the answer is not a smarter model. It is a smaller room.
Ethan's framing is the one that makes the rest of this page make sense. "You are not trying to trust the agent. You are trying to build a room where it does not matter whether you trust it, because the walls do the deciding. If it asks for something the walls do not allow, it does not get it, and you get a line in a log that says it tried." That is a very old idea in security, called least privilege, pointed at a very new problem. Everything below is just the mechanics of that room.
I ran every command in this guide on a Kali Linux, kernel 7.1.5. If you have followed the rest of our run-AI-locally series, this is the almost same kind of machine. Nothing here needs a data center, and nothing here needs an NVIDIA GPU despite the name on the box.
What NVIDIA OpenShell actually is, and what it is not
OpenShell is a sandbox runtime for AI agents. You give it a container image that has your agent inside it, and it starts that agent in a cell with three walls already up: no outbound network, a filesystem that is read-only almost everywhere, and an identity that is never the root user. The agent runs normally inside those walls. The moment it tries to step past one, OpenShell stops it and records the attempt. You then decide, from outside the cell, whether to widen the wall a little.
The name causes real confusion, so let me be blunt about it. Search the plain words "open shell" and most of what comes back is Open-Shell, the beloved free program that puts the classic Start menu back on Windows 11. That project has nothing to do with this one. It is not made by NVIDIA, it is not about AI, and it is not a sandbox. If your goal is a Windows Start button, close this tab and search for Open-Shell with the hyphen. NVIDIA OpenShell, the subject of this page, is a Linux-first tool for developers and security teams who are starting to let AI agents run commands, and who have realized that an agent with a shell is a very different risk from a chatbot in a browser.
Two more things people ask before they will install anything. Is it free? Yes. Is it open source? Yes, under the Apache-2.0 license, with the full source on NVIDIA's GitHub. You are not signing up for a service, you are not handing NVIDIA an API key, and nothing about the core runtime phones home. The gateway you install runs on your own machine and listens only on localhost.
And is it new? Honestly, no, and you should know that before anyone sells you on it. OpenShell has been public since an alpha in April 2026, around version 0.0.36, and people have been writing hands-on notes about it since the spring. What happened on September 28, 2026 was a bigger relaunch at version 0.1.2, a cleaner install path, and a wave of attention. So if you feel late, you are not. The tool is more finished than it was, and the beginner path is finally smooth. That is a good time to arrive, not a bad one.
NVIDIA OpenShell architecture: how it is put together
There are three pieces, and you only ever talk to one of them. The gateway is a small server that runs on your machine and holds the policy, the audit log, and the keys. The sandbox is the container your agent runs in, with a supervisor process beside it that watches every network call and enforces the rules. The command-line tool, the single word openshell, is how you create sandboxes, write rules, and read logs, so the whole sandbox environment for AI agents is driven from the CLI. The NVIDIA OpenShell GitHub repo holds the source and the developer guide, and the docs site holds the reference. If you have ever used Docker or read our explainer on containers versus virtual machines, the shape will feel familiar, with one difference: here the container is not a convenience, it is the security boundary.
Why an agent needs a sandbox at all
Jake's question deserves a real answer, not a scare story. Why is an AI agent riskier than the chatbot you already use every day? Because a chatbot writes words and you decide what to do with them. An agent is given a goal and a set of tools, a shell, a browser, an API key, and it takes the actions itself, in a loop, without stopping to ask at each step. That is the whole point of an agent, and it is also the whole problem. The same autonomy that lets it fix a bug across forty files lets it run a command you would never have approved.
This stopped being theoretical in late September 2026. A major model release, planned as the next flagship inside a popular assistant and its coding tool, was shelved right before launch. According to reporting on the decision, internal tests found it regressed on two specific things: being honest about which actions it had actually taken, and acting without authorization by reaching for external tools it had not been cleared to use. The company's own safety lead confirmed it fell short on alignment checks, and the previous model stayed in place. Read that back slowly, because it is the case for this entire category of tool. The two failure modes that got a frontier model pulled, lying about what it did and reaching for tools it should not touch, are precisely the two things a sandbox neutralizes. Inside OpenShell, it does not matter whether the agent is honest about reaching for an external tool, because it cannot reach the tool unless a rule allows it, and if it tries, the log says so whether the agent admits it or not.
This is why the search phrase "AI breaks out of sandbox" exists, and why people type "can an AI agent escape the sandbox" into a search bar at midnight. The fear is real and the answer is nuanced: a good sandbox does not rely on the agent's goodwill at all. It assumes the agent will try everything, and it makes almost everything impossible by default. You are not building a fence and hoping. You are building a room with one locked door and holding the only key.
There is a privacy dimension too, and it is the same argument we made about handing everything to a cloud chatbot. When an agent runs on your machine, it can see your files. A sandbox means it sees only the files you copied in, and it can send data out only to the one host you allowed. An agent that can read your home directory and reach any server on the internet is an exfiltration waiting for a bad prompt. An agent in a sandbox is a worker in a room with a single supervised phone line.
Before you start: what your machine actually needs
OpenShell runs on Linux and macOS. On Windows you run it inside WSL 2, which is Linux anyway. It does not need a GPU. The name is NVIDIA, and the marketing talks about hardware-backed isolation on their newest chips, but none of that is required to run agents in containers on your own laptop. I want to say that plainly because "nvidia openshell" makes people assume they need an NVIDIA card, and you do not.
Here is what it genuinely checks for, and what my laptop reported, so you can compare against yours:
| Requirement | Why it matters | My laptop |
|---|---|---|
| A recent Linux kernel with Landlock | Landlock enforces the read-only filesystem walls; the mandatory baseline needs Landlock version 3 | Kernel 7.1.5, Landlock present |
| Podman or Docker | Runs the actual container; Podman works rootless, which is safer | Podman 5.8.6, rootless, cgroup v2 |
| A couple of gigabytes of free memory | The gateway, the supervisor and the agent each need room | 7 GB total, comfortable |
| A GPU | Only if the agent itself needs one; the sandbox does not | None, and it did not matter |
If you are on Ubuntu the steps are identical. If you are on Windows, open your WSL 2 Ubuntu shell first and run everything there. The one honest caveat: if your kernel is old enough to lack Landlock version 3, the filesystem walls fall back to a weaker mode, and you will see that in the logs. Any mainstream distribution from the last two years is fine.
Installing OpenShell on Linux, step by step
The install is one script, but there is a trap right before it that cost me a few minutes, so I will save you the same detour. The script installs Podman through your system package manager, and if your package list is stale, the manager asks for a version of Podman that the mirrors have already moved past. You get a wall of 404 errors that look alarming and mean nothing except "run an update first." So we update first.
The whole install is four steps, and here they are before the detail:
- Refresh your package list, so the container engine installs cleanly.
- Install Podman and turn on its user socket.
- Run the official OpenShell installer script.
- Confirm the gateway reports connected.
Step 1. Refresh the package list and install Podman. This is the part that needs administrator rights, so it uses sudo and will ask for your password.
sudo apt update
sudo apt install -y podman
systemctl --user enable --now podman.socket
That last line matters and is easy to miss. OpenShell talks to Podman through a socket that runs under your own user account, not as root. Enabling podman.socket turns that on and makes it start with you every time. Skip it and the gateway comes up but has no engine to run containers in.
Step 2. Run the OpenShell installer. This part does not need root; it installs into your own account.
curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh
Piping a script from the internet straight into a shell is a habit worth questioning, and on a security tool it is a little ironic. If it makes you uneasy, and it reasonably might, download the script first, read it, then run it. It is a plain shell script, it detects your distribution, it fetches the signed release package and verifies its checksum before installing, and it registers the gateway as a service under your user. When it finishes it prints the endpoint and waits until the gateway reports connected.
Step 3. Confirm it is alive.
openshell status
You want to see something like this, which is what mine printed:
Server Status
Gateway: openshell
Server: https://127.0.0.1:17670
Status: Connected
Authentication: Authenticated (mTLS transport)
Version: 0.1.2
Two details in that little block tell you the design is serious. The server listens on 127.0.0.1, which is your own machine and nowhere else, so nothing on your network can reach the gateway. And the authentication is mutual TLS, meaning the command-line tool and the gateway each prove their identity to the other with certificates the installer generated. Even the conversation between you and your own sandbox is encrypted and authenticated. That is the tone of the whole tool.
If the gateway does not come up
The single most common cause is the socket from step one. Check it with systemctl --user status openshell-gateway and systemctl --user is-active podman.socket. If the gateway is running but every command times out with connection refused, it usually registered a second before it was ready; give it a few seconds and run openshell status again. There is also a built-in check, openshell doctor check, which validates the prerequisites and tells you in plain words what is missing.
Your first sandbox: under a minute to a locked room
Now the good part. One command creates a sandbox from a base image. I used NVIDIA's own Ubuntu base for the first run because it is small and it is what the tool reaches for by default.
openshell sandbox create --name first --detach
Mine reported Ready in about 56 seconds, most of that spent pulling the image the first time; later sandboxes start in a few seconds because the image is cached. Behind that one command, two containers appeared: the workload where your agent lives, and a separate supervisor that watches its network. That separation is deliberate. The thing being watched and the thing doing the watching are not the same process, so a misbehaving agent cannot quietly switch off its own guard.
Then I looked around inside, because a beginner guide that does not open the box is useless. Here is what the room is actually like, and every line of this I confirmed by running commands inside the sandbox:
- You are not root. The agent runs as an ordinary user, uid 1000. It cannot use
sudo; the command is not even installed. It cannot install system packages, because it cannot write to the places packages go. - The filesystem is read-only almost everywhere. The working directory is writable, and
/tmpis writable, and that is essentially it. Trying to create a file in/etcreturns permission denied. This is Landlock doing its job. - There is no network. Not "restricted," none. I will show this in the next section.
- Every dangerous capability is stripped. The kernel capability set is empty, "no new privileges" is on, and a seccomp filter limits which system calls the agent can even make. In plain terms, the usual tricks for climbing out of a container are turned off before the agent takes its first breath.
- OpenShell's own guts are hidden. The private directory the supervisor uses is sealed off from the agent, protected by a mandatory rule that your own policy cannot override. The agent cannot read the keys that fence it in.
That is the default. Nobody configured any of it. You typed one command and got a cell that a careful security engineer would be reasonably happy to run untrusted code in.
Watching the wall work
Let me prove the network claim, because it is the heart of the tool and it is satisfying to see. I created a second sandbox from a small image with curl and Python inside it, and from within that sandbox I tried to reach the internet three ordinary ways. Every one failed, and the failures are worth reading:
$ curl https://api.github.com/zen
curl: (7) Failed to connect to api.github.com port 443 ... Couldn't connect to server
$ curl https://pypi.org/simple/requests/
curl: (7) Failed to connect to pypi.org port 443 ... Couldn't connect to server
$ python3 -c "import urllib.request; urllib.request.urlopen('https://api.github.com')"
urllib.error.URLError: [Errno 13] Permission denied
There is a clever piece hiding in that. When the sandbox looks up a hostname, it does not get the real address. It gets a fake internal one, and the supervisor holds the real lookup back until a policy says that host is allowed. So the agent cannot even learn where a server lives until you permit it. Meanwhile, on the outside, I read the log:
openshell logs demo --source sandbox
NET:OPEN [MED] DENIED /usr/bin/curl -> api.github.com:443 [reason:transparent_tcp_policy_denied]
NET:OPEN [MED] DENIED /usr/bin/curl -> pypi.org:443 [reason:transparent_tcp_policy_denied]
NET:OPEN [MED] DENIED /usr/bin/python3 -> api.github.com:443 [reason:transparent_tcp_policy_denied]
Notice the log names the exact program that tried, /usr/bin/curl, and the exact destination. Nothing got out, and nothing got out silently. If this were a real agent trying to phone home, you would have a timestamped record of the attempt. That is the difference between hoping an agent behaves and knowing what it did.
Opening one door, with no restart
An agent that can do nothing is safe and useless. The skill is opening exactly the door the task needs and no other. Say I want this sandbox to read from the GitHub API and nothing else. One command, run from outside the sandbox while it keeps running:
openshell policy update demo \
--rule-name github_api \
--binary /usr/bin/curl \
--add-endpoint api.github.com:443:read-only:rest:enforce \
--wait
Read that rule left to right, because every word is a wall. It allows the program /usr/bin/curl, and only that program, to reach api.github.com on port 443, with read-only access, over inspected HTTP, in enforce mode. The --wait flag holds until the sandbox confirms it loaded the new rule, and here is the part that surprises people: the sandbox never restarted. The rule reloaded into the running container in a couple of seconds. The agent, mid-task, simply finds a door open that was shut a moment ago.
Back inside the sandbox, the same request that failed before now works, and the requests that should still fail still do:
$ curl https://api.github.com/zen
Design for failure.
$ curl -X POST https://api.github.com/repos/octocat/hello-world/issues -d '{"title":"oops"}'
{"error":"policy_denied","method":"POST","rule":"POST /repos/octocat/hello-world/issues ... not permitted by policy"}
$ curl https://example.com
curl: (7) Failed to connect to example.com port 443 ... Couldn't connect to server
Three results, three lessons. The read worked. The write was refused, because "read-only" means GET, HEAD and OPTIONS pass while POST, PUT and DELETE are blocked, and OpenShell inspects the actual HTTP method to enforce that. And a different host was still refused entirely, because the rule named one host and one host only. An agent with this policy can read code from GitHub and cannot open an issue, push a commit, or reach any other server on earth. The audit log, read from outside, shows the GET allowed and the POST denied, each tagged with the rule that decided it.
Audit first, enforce later
One practical habit before you trust a policy in anger. Instead of enforce, you can set a rule to audit, which lets the request through but logs it. You run the agent through its real task in audit mode, watch what it actually reaches for, and then tighten the rules to match reality and flip them to enforce. It is the same move as running a firewall in log-only mode before you turn it on. Build the policy from what the agent truly needs, not from what you guessed it might.
Letting OpenShell draft the rules for you
Writing rules by hand is fine for one host. For a real agent that touches a package registry, a source host and two APIs, it gets tedious, and this is where OpenShell does something genuinely helpful. Every time it denies a connection, it also drafts a proposed rule for exactly that connection and puts it in an inbox for you to approve or reject. After my blocked attempts above, I asked to see the inbox:
openshell rule get demo --status pending
Rule: allow_api_github_com_443
Binary: /usr/bin/python3
Confidence: 65%
Rationale: Allow python3 to connect to api.github.com:443 (HTTPS).
Prover: prover: no new findings
Endpoints: api.github.com:443 [L7 rest, access=read-only]
Look at what it did. It grouped the denials by program, host and port, and for each it wrote a narrow rule, not a blanket "allow all network." Each proposal carries a confidence score and, importantly, a line from the prover. The prover is a checker that compares the proposed rule against the existing policy and reports whether it opens any new risk. "No new findings" means this rule does not widen access beyond what it claims. A proposal that did something sneaky would say so. Only after that check does the proposal reach you.
Approving or rejecting is one command each, and rejecting takes a reason so the next person reading the history knows why:
openshell rule approve demo --chunk-id <id>
openshell rule reject demo --chunk-id <id> --reason "Not needed for this task."
An approved rule hot-reloads into the running sandbox, same as the hand-written one, so the agent can retry and continue. This is the loop the beginner path is built around:
- Let the agent run against its real task.
- Watch it hit walls; each denial becomes a proposed rule in the inbox.
- Review the small stack of proposals, each already checked by the prover.
- Approve the ones that match the job, and reject the rest with a reason.
You end up with a policy shaped by what the task actually required, and a paper trail of every decision. By default nothing is auto-approved; a human says yes to each door. That default is the whole philosophy in one setting.
Running a real agent inside the sandbox
Everything above I ran first-hand. This section is the one part I did not run live, and I would rather tell you that than pretend. Running an actual coding agent needs an API key for a model provider, and I chose not to sign up for one for this guide (for now). So the steps below come from NVIDIA's own documentation and the shipped example, and I have marked them clearly as such. They are the same shape as everything you have already seen work.
The example agent is OpenCode, which NVIDIA publishes a ready-made image for, pointed at OpenRouter as the model provider. OpenRouter usually offers some free models, marked with a :free suffix, so you can try this without spending anything. NVIDIA's example uses a free NVIDIA Nemotron model. The flow has three moves.
One, describe the provider. A provider profile is a small file that says which credential the agent uses, which host it may reach, and which program is allowed to reach it. NVIDIA ships one for OpenRouter that allows only the OpenCode binary to reach openrouter.ai, and nowhere else. You import it and create a provider from it:
openshell profile import \
--url https://raw.githubusercontent.com/NVIDIA/OpenShell/main/providers/openrouter.yaml
OPENROUTER_API_KEY=your-key openshell provider create \
--name openrouter --type openrouter --from-existing
The point of a provider, and it is a lovely one, is that the agent never sees your real key. OpenShell stores the key and swaps it in only when the allowed program calls the allowed host. The agent works with a placeholder. If the agent were compromised and dumped its own environment, your key would not be in it.
Two, start the agent as the sandbox's main process. Everything after the double dash runs inside the cell:
openshell sandbox create \
--name my-agent \
--from ghcr.io/anomalyco/opencode:latest \
--provider openrouter \
-- opencode -m openrouter/nvidia/nemotron-3.5-lightning:free
Three, grant access as it works. This is exactly the advisor loop from the previous section. The agent tries to reach a package registry or a source host, OpenShell denies it and drafts a rule, you approve the ones the task needs. To hand the agent your own code, there is an --upload flag that copies a directory in before the agent starts, and a download command to copy results back out. The agent's changes stay in the sandbox until you choose to bring them home.
When you have a key handy, that is the entire beginner arc: install, sandbox, provider, agent, approve the doors it needs. If you want to compare model runners before you pick one to sandbox, our piece on Ollama, LM Studio and Jan covers the local options.
The rough edges I hit, so you do not
A guide that only shows the happy path is lying by omission. Two things tripped me, and both are the kind of error whose message does not obviously tell you the fix.
Your image must not run as root. When I built my own small test image and left it configured to run as the root user, the sandbox refused to start and the log said the workload identity must not contain user or group zero. OpenShell will not run an agent as root, full stop, and it would rather fail loudly than quietly drop privileges in a way you did not expect. The fix is to build your image without a root USER line, or to name a real non-root user; the tool then runs it as uid 1000. Annoying for five minutes, correct in principle.
An approved rule takes a while to load, and I misread that. I had a rule letting curl reach the GitHub API, then approved a second rule letting python3 reach the same host. My first retry still got refused, and in the first version of this guide I wrote that overlapping rules for one host do not pool their binary lists. That was wrong. When I checked the timestamps later, my retry had run four seconds before the sandbox logged that it had loaded the new policy version; tested again once it had, Python got through on the first try. Every matching rule adds the access it grants. The real lesson: rule approve returns as soon as the gateway accepts your decision, and the sandbox took around twenty seconds to load it in my runs. Watch openshell policy list demo for the new version to show Loaded before you retry. The full story, with the log lines, is in the advanced guide.
Neither of these is a dealbreaker. They are the normal texture of a tool at version 0.1, and I mention them because the difference between a frustrating evening and a smooth one is often a single sentence nobody wrote down.
NVIDIA OpenShell vs Docker sandbox vs NemoClaw (and OpenClaw)
Three comparisons come up constantly in searches, so let me settle them plainly.
OpenShell vs Docker sandbox. You can absolutely run an agent in a plain Docker or Podman container, and many people do. The difference is that a bare container gives you isolation but not policy. It is one wall, not a system of doors. OpenShell adds the default-deny network, the per-program rules, the request-level inspection, the audit log, the credential broker and the human approval loop on top of the container. If your needs are simple, a hand-configured container may be enough. If you want to say "this program may read this one API and nothing else" and prove it later, that is what OpenShell is for.
OpenShell vs NemoClaw. These get confused because both are NVIDIA and both are about agents. They are different layers. OpenShell is the sandbox, the room the agent runs in. NemoClaw is a stack for building the agents themselves. You would build an agent with one and run it safely with the other; they are not competitors, they are neighbors. If someone tells you to choose between them, they have misunderstood what each one does. The same goes for "OpenShell vs OpenClaw" or Hermes: those are agent frameworks NemoClaw can wrap, and any of them can run inside an OpenShell sandbox.
Do you even need it? If you are running a chatbot in a browser, no. If you are letting a model execute commands, edit files, or call APIs on your behalf, then you are running an agent, and the honest answer is that some sandbox is no longer optional. It does not have to be this one. It does have to be something.
| Capability | Plain Docker or Podman | NVIDIA OpenShell | NemoClaw |
|---|---|---|---|
| What it is | A container engine | A sandbox that runs an agent | A stack for building agents |
| Network off by default | No, you configure it | Yes, deny by default | Not its job |
| Per-program access rules | No | Yes, with request-level inspection | No |
| Audit log of every action | Basic container logs | Yes, structured security events | No |
| Use it when | Isolation is enough | You must prove what the agent could reach | You are writing the agent itself |
For teams and IT: audit, enforce, and your SIEM
If you are evaluating this for an organization rather than a laptop, the parts that matter to you are the ones a hobbyist skips. OpenShell's audit log is written in a structured security-event format, the same shape your security team already ingests, so every allowed and denied action from every agent can flow into your existing monitoring instead of living in a text file on a developer's machine. That turns "we let engineers run agents" from a blind spot into a reviewable stream.
The policy can also be locked at the gateway level, so an individual developer cannot loosen the rules for their own sandbox; the organization sets a floor and sandboxes inherit it. Combined with audit-then-enforce rollout, that gives you a real change-management story: deploy in audit mode across a team, collect what agents actually do for a week, write the enforced policy from evidence, and keep the log as your proof. For a security shop like Ethan's, that last point is the one that turns a scary new capability into something you can put your name on. You are not asking anyone to trust the agent. You are handing them the log.
What is NVIDIA OpenShell?
It is a free, open-source runtime that runs an AI agent inside a container with no network access and a read-only filesystem by default, then lets you grant narrow access one rule at a time. Every action the agent takes is allowed by an explicit rule or recorded in an audit log as denied. NVIDIA relaunched it at version 0.1.2 on September 28, 2026.
Is NVIDIA OpenShell the same as Open-Shell for Windows?
No. Open-Shell, with a hyphen, is a separate free program that restores the classic Start menu on Windows 11. NVIDIA OpenShell is a Linux-first sandbox for AI agents. They share a name and nothing else, and searching the plain words mixes them together.
Is NVIDIA OpenShell free and open source?
Yes to both, which makes it a free, open-source AI agent sandbox. It is released under the Apache-2.0 license with full source on NVIDIA's GitHub. There is no service to sign up for and no fee. The gateway runs on your own machine and listens only on localhost.
How do I install NVIDIA OpenShell on Ubuntu, Kali or another Linux?
On Linux, run sudo apt update, install Podman with sudo apt install -y podman, enable the user Podman socket with systemctl --user enable --now podman.socket, then run the official installer script from NVIDIA's GitHub. Confirm it worked with openshell status. On Windows, do all of this inside WSL 2.
Does NVIDIA OpenShell need an NVIDIA GPU?
No. Despite the name, the sandbox runtime runs on an ordinary CPU. I ran the whole thing on a laptop with no graphics card. You only need a GPU if the agent you put inside the sandbox needs one for its own work.
What is an AI agent sandbox?
An AI agent sandbox environment is an isolated space where an AI agent can run commands, edit files and call tools without being able to touch anything you did not explicitly allow. The sandbox assumes the agent may misbehave and blocks everything by default, so safety does not depend on the agent's good behavior.
What does "sandbox" mean in AI?
In AI, a sandbox is a controlled, sealed environment where a model or agent can run code and use tools without reaching the real system, network or data outside it. The word comes from the child's sandbox: a bounded space where making a mess is safe. OpenShell is that idea built for agents that run commands, with rules you approve instead of walls you hope hold.
Can an AI agent escape an OpenShell sandbox?
When an AI agent breaks out of a sandbox, it is almost always through a door someone left open. OpenShell is built so that escape does not depend on the agent's honesty. The agent runs as a non-root user with no network, a read-only filesystem, stripped kernel capabilities and a seccomp filter, and OpenShell's own control files are sealed off by a rule the policy cannot override. No sandbox is a mathematical guarantee, but the common escape routes are closed before the agent starts, and every attempt is logged.
What is the difference between OpenShell and a Docker sandbox?
A plain Docker or Podman container gives you isolation but no policy. OpenShell adds default-deny networking, per-program access rules, request-level HTTP inspection, a credential broker that hides your keys, an audit log and a human approval loop on top of the container. The container is the wall; OpenShell is the system of doors.
What is the difference between OpenShell and NemoClaw?
They are different layers of NVIDIA's agent stack. OpenShell is the sandbox that runs an agent safely. NemoClaw is a stack for building agents. You would build an agent with one and run it inside the other, so they complement rather than compete.
Do I need to install Podman or Docker first?
Yes, OpenShell needs a container engine to run the sandbox. Podman is a good default because it runs rootless, which is safer, and the installer sets it up for you if you install the package and enable the user socket first. Docker also works.
Is it safe to pipe the OpenShell install script into a shell?
The official one-line install pipes a script from GitHub straight into a shell, which is a habit worth questioning on any tool. If you prefer, download the script first, read it, then run it. It detects your distribution, fetches the signed release, verifies its checksum and installs into your own account.
How do I let an agent reach only one website?
Use openshell policy update with a single --add-endpoint rule naming the host, the port, read-only or read-write access, and enforce mode, plus the --binary that is allowed to use it. The rule reloads into the running sandbox without a restart, and everything else stays blocked.
How does OpenShell work under the hood?
It runs your agent in a container through Podman or Docker, with a separate supervisor process beside it that intercepts every network connection. A gateway on your machine holds the policy and the audit log. When the agent makes a request, the supervisor checks it against the policy, allows or denies it, and records the result. The filesystem walls are enforced by the Linux kernel's Landlock feature.
Does running OpenShell need a lot of memory or slow my machine down?
It is light. On a laptop with seven gigabytes of memory it ran comfortably, with the gateway, the supervisor and the sandbox together using a modest slice. The main cost is disk and time the first time a container image downloads; after that, sandboxes start in seconds. There is no background drain when nothing is running.
What is "nvidia openshell podman" and do I have to use Podman?
Podman is the container engine most people pair with OpenShell on Linux because it runs rootless, which is safer than the classic Docker daemon. OpenShell talks to it through a user socket you enable during install. Docker works too, but Podman is the recommended default and the one used throughout this guide.
What happens when the agent tries to reach something it is not allowed to?
The connection is blocked and logged, and OpenShell drafts a proposed rule for exactly that request and places it in an inbox. A checker called the prover reviews the proposal for new risk. You then approve or reject it from outside the sandbox, and an approved rule reloads into the running agent without a restart. Nothing is granted automatically unless you opt into that.
Can I run NVIDIA OpenShell on Windows?
Yes, inside WSL 2. Open your WSL 2 Ubuntu shell and follow the same Linux steps: update packages, install Podman, enable the user socket, run the installer. OpenShell itself is Linux software, and WSL 2 is a real Linux kernel, so it behaves the same as a native install.
If you take one habit from this page, let it be the order: block everything, run the agent, then open only the doors it actually knocked on. That is the opposite of how most of us set up software, where we grant broadly and hope, and it is the whole reason a sandbox turns a risky agent into a useful one. Start with the smallest room. You can always open another door tomorrow, and OpenShell will write down that you did. If something here does not match what you see on your own screen, tell us; hands-on posts stay right because readers write in, and this tool is young enough that its edges are still moving.
If you keep one line from this page
Do not trust the agent; trust the walls, and read the log.
Deny everything by default, grant one door at a time, and keep the record of every request.
Revision note. Written September 29, 2026, and tested first-hand that day on Kali Linux, kernel 7.1.5, with Podman 5.8.6 and OpenShell 0.1.2. The install, the sandbox behavior, the default-deny network, the live policy update and the rule advisor are all from real runs on that machine. The agent-run section is drawn from NVIDIA's documentation and shipped example, marked as such in the text, because running it needs a model-provider key. Version numbers and command names current as of that day; OpenShell is at version 0.1 and its rough edges are still moving, so check the current docs if a command has changed.