NVIDIA OpenShell Advanced: Policy YAML, Rules, Prover, MCP

Logeshwaran
—

This is the advanced NVIDIA OpenShell guide: the full OpenShell policy YAML, per-binary OpenShell network policy rules that inspect paths and methods, audit mode before enforce mode, the policy advisor and its approval modes, the policy prover, credentials, MCP rules, the DNS exfiltration hole that was closed in May, filesystem and process limits, global policy for a fleet, and every rough edge I hit. It follows our OpenShell beginner guide, which installs NVIDIA's code sandbox for AI agents and opens one door. Everything below that says "I ran" was run on the same Kali Linux laptop, OpenShell 0.1.2 on Podman 5.8.6, on September 29, 2026, and the log lines are pasted as the tool printed them. The parts I could not run, MCP and credentialed providers, are marked as such. One correction is in here too, to something I wrote in the beginner post, because a second look proved me wrong.

Jake read the beginner guide and came back with the question every team asks next. "Fine, one door for one program. But we want an agent working over a client's repo for a week. I am not going to sit at a rule inbox all day, and if I turn on auto-approve, what am I actually approving?" Ethan, who runs the shop, wanted the harder version of that question answered before anyone said the word "client." So I went and found out. The short answer surprised me, and it is the first thing on this page.

⚡ Quick Answer

• The surprise → with auto-approval on, the advisor approved a connection-only rule for curl, and a POST to GitHub went straight through. A hand-written read-only rule would have blocked it. What happened.

• Policy YAML → five sections; filesystem, Landlock and process are fixed at start, network rules and middleware reload live. Line by line.

• Rules that read the request → allow GET /repos/**, deny /repos/*/*/rulesets, and deny always wins. Real output.

• Prove it → openshell-prover check candidate.yaml --boundary boundary.yaml returns within, exceeds, or unsupported, with an exit code for CI. How.

The DNS exfiltration issue from May 2026 is covered from the issue and the fix, not by running anything. Section.

If the beginner guide was about the room, this one is about the locks. Nothing here needs a GPU, a cloud account, or a model key, except the two sections that say so. If you are new to agents entirely, our plain-words piece on what an AI agent is comes first, then the beginner guide, then this.

How an OpenShell policy is put together: the five sections

A sandbox policy is one YAML file with version: 1 at the top and up to five sections under it. The single most useful thing to know about them is which ones can change while the sandbox is running, because that decides how you work with the tool day to day.

Section What it controls Enforced by Changes live?
filesystem_policyWhich paths are readable, which are writableLinux Landlock, at startNo, fixed at start
landlockWhether the sandbox may start if Landlock cannot be appliedSupervisor, at startNo
processThe user and group the agent runs asDocker or Podman, at createNo
network_policiesWhich program may reach which host, port, path and methodSupervisor proxy, OPA engine, L7 inspectorYes, hot reload
network_middlewaresInspect, redact or block requests and responses in flightSupervisor, after policy, before credentialsYes

I proved the "fixed at start" claim rather than take it on trust, because the tool will let you believe otherwise. I pushed a new policy to a running sandbox that added /var/tmp to the writable list. The gateway accepted it, reported version 2 loaded, and openshell policy get --base dutifully listed /var/tmp as writable. Inside the sandbox, touch /var/tmp/x still said permission denied. The filesystem walls were built when the container started and they do not move. That is correct behavior, and also a trap: the policy you can read is not always the policy being enforced, for these three sections. If you need a filesystem or user change, stop the sandbox and recreate it with the new file.

The NVIDIA OpenShell docs and developer guide on GitHub describe each field; the NVIDIA OpenShell architecture behind them is the one from the beginner guide, a gateway, a supervisor beside each sandbox, and the command-line tool. Two more facts about the shape. There is a mandatory baseline underneath your policy that you cannot loosen: the private /.openshell directory, where the supervisor keeps its channel and keys, is sealed by a Landlock rule that requires Landlock ABI version 3, and the best_effort compatibility setting does not relax it. And when a policy file is present but a section is missing, the defaults are not what you might guess, which is the subject of the first gotcha below.

The full policy YAML, line by line

Here is the policy I ended up running on my demo sandbox. It is small enough to read in one sitting and it exercises the interesting parts: two binaries in one rule, path-level allow rules, a deny rule, and one endpoint in audit mode.

version: 1
filesystem_policy:
  include_workdir: true
  read_only: [/bin, /usr, /lib, /proc, /dev/urandom, /etc, /var/log]
  read_write: [/tmp, /dev/null]
landlock:
  compatibility: best_effort
network_policies:
  github_api:
    endpoints:
      - host: api.github.com
        port: 443
        protocol: rest
        enforcement: enforce
        rules:
          - allow:
              method: GET
              path: /zen
          - allow:
              method: GET
              path: /repos/**
        deny_rules:
          - method: "*"
            path: "/repos/*/*/rulesets"
    binaries:
      - path: /usr/bin/curl
      - path: /usr/bin/python3.12
  pypi_audit:
    endpoints:
      - host: pypi.org
        port: 443
        protocol: rest
        enforcement: audit
        access: read-only
    binaries:
      - path: /usr/bin/curl

Reading it top to bottom. The filesystem block is the built-in default written out, and I recommend always writing it out, for a reason coming in a moment. include_workdir: true makes the working directory, /sandbox in my image, writable; it defaults to true when there is no policy file at all and to false the moment you supply one. The seven read-only paths are what a normal Linux program needs to run. Everything not listed is invisible: inside my sandbox, ls / itself returns permission denied, and /home does not exist as far as the agent can tell.

The network block is a map of named rules. Each rule has endpoints, the destinations it allows, and binaries, the programs it allows to reach them. A rule with an endpoint in it is only ever an allow; the only way to say no is a deny_rules entry inside an inspected endpoint, or simply not writing a rule. The github_api rule uses explicit rules instead of an access preset, which is what lets it say "only these two paths." You cannot combine access and rules on one endpoint; pick one. The pypi_audit rule uses the preset read-only, which means GET, HEAD and OPTIONS, and sets enforcement: audit, which we will come back to.

Set, update, dry-run and versions

There are two ways to change the network part of a live policy, and they are not interchangeable. openshell policy update merges small changes into what is there: add an endpoint, add an allow or deny line to a named rule, remove a rule. openshell policy set --policy file.yaml replaces the whole policy with your file. Anything that update cannot express, and that includes middleware, GraphQL, MCP and JSON-RPC rules, TLS skipping, credential bindings and request signing, has to go through set. I applied the file above with:

openshell policy set demo --policy policy-advanced.yaml --wait
✓ Policy version 4 submitted (hash: 2d553a746951)
✓ Policy version 4 loaded (active version: 4)

Every change gets a version number and a hash, and openshell policy list demo shows the history with each one marked Loaded, Superseded or Pending, plus an error column for a revision the sandbox refused. You can read any old revision with --rev. This is your change record, and it is one of the reasons the tool is defensible in front of an auditor: the policy in force at 12:44 on a given day is a hash you can quote.

Before you send an update, there is a flag I wish every config tool had. --dry-run prints the merged policy that would result, without sending it. I used it to see what adding an audit-mode pypi rule would do to the existing rules, and it printed the full YAML with the new rule slotted in and nothing else touched. Cheap insurance.

One more distinction that matters once you attach credentials. policy get --base shows the policy you wrote. policy get --full shows the effective policy, which is yours plus any rules contributed by a provider you attached, such as a GitHub provider quietly adding the GitHub API endpoint for its own binary. When something reaches a host you never wrote a rule for, look at --full before you assume a bug.

Gotcha: a policy file replaces the default entirely

This one cost me two sandboxes. I wanted to test the process section, so I wrote a tiny file with just version: 1 and process: {run_as_user: "1500", run_as_group: "1500"} and passed it with --policy at create time. The sandbox went to phase Error within seconds, "Container exited with code 1", and the workload log said the sandbox boundary received a termination signal before the supervisor confirmed. It was not the user ID. A policy file does not layer on top of the default; it is the policy. With no filesystem_policy block, the agent had no readable paths at all and could not even start. The same file with the filesystem block copied in from above worked immediately: id inside printed uid 1500, the working directory was writable, and whoami complained it could not find a name for user 1500, which is harmless. Always start a custom policy from the full default, never from a fragment.

Network rules that read the request: paths, methods and deny rules

The beginner guide opened one host with a read-only preset. The advanced version is a rule that reads the actual HTTP request and decides per path and method. This is where OpenShell stops being a firewall and becomes something closer to an API gateway that happens to live next to your agent, and it is where AI agent networking security stops being a diagram and becomes a rule you can read.

The endpoint syntax on the command line is host:port[:access][:protocol][:enforcement], and the same fields appear as keys in YAML. Three of them deserve a table, because the choices are not obvious from the names.

protocol What OpenShell inspects Set with
restHTTP method, path, query parametersupdate or set
websocketThe upgrade request and client text frames; not binary frames, not server messagesupdate or set
graphqlOperation type, operation name, top-level fieldsset only
mcpMCP method and tool name in the request bodyset only
json-rpcJSON-RPC method nameset only
tcpConnection only; nothing inside itupdate or set
omittedNo request rules; TLS is still terminatedupdate or set

The access presets: read-only is GET, HEAD and OPTIONS; read-write adds POST, PUT and PATCH but not DELETE; full is everything. The enforcement field is enforce or audit, and the default, if you leave it out, is audit. That default matters and I will come back to it.

Now the part worth the price of admission. With the policy above loaded, I ran four requests from inside the sandbox. Here they are with what came back, unedited apart from trimming long JSON:

$ curl https://api.github.com/zen
Responsive is better than fast.

$ curl https://api.github.com/users/octocat
{"error":"policy_denied","layer":"l7","method":"GET","path":"/users/octocat",
 "detail":"GET /users/octocat not permitted by policy","policy":"github_api"}

$ curl https://api.github.com/repos/octocat/hello-world
{ "id": 1296269, "name": "Hello-World", ...

$ curl -w "%{http_code}" https://api.github.com/repos/octocat/hello-world/rulesets
403  {"error":"policy_denied","detail":"GET /repos/octocat/hello-world/rulesets blocked by deny rule"}

Read those as a set. /zen matched an allow line. /users/octocat matched nothing, so it was refused even though the host and method were fine. /repos/octocat/hello-world matched the /repos/** glob. And the rulesets path matched both the /repos/** allow and the deny rule, and the deny won, with the log saying so in words: "blocked by deny rule." That last one is the rule to remember about how rules combine. Network rules are not an ordered firewall list where the first match decides. Every matching allow adds what it allows, and a single matching deny beats all of them, wherever it sits. The agent got a clean JSON error body it can read, which is deliberate: a well-behaved agent sees policy_denied and the exact path, and can ask for a rule instead of retrying blindly.

The glob rules are worth two sentences because they are stricter than shell globs. In a path, * matches within one segment and ** matches across slashes, but only as a whole segment, so /repos/** is right and /repos/**x quietly behaves like a single star. In a host, the same idea applies to dots: *.internal.example matches one label, a wildcard never spans dots, a bare * is rejected, and wildcards need at least three labels. The validator refuses a policy that breaks these, and it refuses at load time, which is the right time.

Adding a single path rule from the command line, without rewriting the file, looks like this. Note that for these L7 appends you must name the rule and either list every binary again or say --any-binary; the tool will not guess:

openshell policy update demo --rule-name github_api \
  --binary /usr/bin/curl --binary /usr/bin/python3.12 \
  --add-allow 'api.github.com:443:GET:/users/*' --wait

Per-binary rules and binary identity

Every rule names the programs it applies to, and this is the feature that separates OpenShell from a network firewall. A firewall knows about addresses. OpenShell knows that /usr/bin/curl asked, and that /usr/bin/python3.12 asked, and treats them differently. I wanted to know how literal that identity is, so I tried the two obvious tricks an agent might stumble into.

First, a symlink. Inside the sandbox I linked /tmp/mycurl to /usr/bin/curl and called the link. It worked, and the log recorded the caller as /usr/bin/curl. The rule matches the real path, resolved by the kernel from the running process, not the name you typed. Second, a copy. I copied the curl binary to /tmp/curl-copy and called that. Refused, exit 7, and the log said exactly why:

NET:OPEN [MED] DENIED /tmp/curl-copy(0) -> api.github.com:443 [reason:transparent_tcp_policy_denied]

Same bytes, different path, no access. That is the behavior you want: an agent cannot borrow a permission by copying a permitted tool somewhere else, and it cannot lose one by reaching a tool through a link. Two further details from the docs and the logs. Rules also match child processes of a listed binary, so a permitted shell script's curl call is judged by the script's interpreter chain, and the denial message helpfully prints the ancestors it saw, with a "SYMLINK HINT" explaining that the identity is the kernel-resolved target. And on the first connection from a binary, its hash is verified, so a permitted path with different contents is caught. If you are unsure what path to write in a rule, run readlink -f $(which python3) inside the sandbox; on my Ubuntu image that is /usr/bin/python3.12, not /usr/bin/python3, and the rule must say the former.

The correction: two rules for one host do merge

In the beginner guide I wrote that when I approved a second rule letting Python reach the GitHub API, alongside an existing rule for curl, Python stayed blocked, and I concluded that overlapping rules for one host do not pool their binary lists. That was wrong, and I want to be precise about why, because the real lesson is more useful than the false one. When I went back through the timestamps, my Python retry ran four seconds before the sandbox logged "Policy reloaded successfully" for the new version. I had tested against the old policy. This evening, with version 3 loaded, the same Python request returned a GitHub aphorism on the first try. Every matching rule adds the access it grants, exactly as the docs say. The true gotcha is a timing one: openshell rule approve returns as soon as the gateway accepts the decision, not when the sandbox has loaded it, and the reload took about twenty seconds in my runs. Watch openshell policy list for the new version to flip from Pending to Loaded, or grep the log for CONFIG:LOADED, before you retry. The beginner post has been corrected to say this.

Audit first, then enforce: building a policy from evidence

Here is a habit that turns policy writing from guesswork into observation. Set a rule to enforcement: audit, let the agent do its real job, and read what it reached for. Then tighten the rule to match and flip it to enforce. It is the same move as running a firewall in log-only mode for a week before turning it on, and it is the reason the default enforcement, when you forget to specify one, is audit rather than enforce: the tool would rather you see traffic you did not expect than block traffic you needed and not know why.

Audit mode did what it promised in my run, with one honest caveat about how it reads. My pypi_audit rule was read-only in audit mode. A GET to the package index returned 200, as expected. A POST, which read-only would block, returned 405 from PyPI itself, meaning the request went all the way through. Both were logged. But look at how the violating POST was logged:

HTTP:GET  [INFO] ALLOWED GET  http://pypi.org:443/simple/requests/ [policy:pypi_audit engine:l7]
HTTP:POST [INFO] ALLOWED POST http://pypi.org:443/simple/requests/ [policy:pypi_audit engine:l7]

Both lines say ALLOWED, and the second does not carry a marker saying "this would have been denied in enforce mode." The policy name in brackets tells you which rule let it through, and you know that rule is read-only, so you can infer the violation, but the log does not infer it for you. So when you review an audit run, do not grep for DENIED; you will find nothing. List every request that passed under an audit-mode rule, and compare each against what the rule would allow in enforce mode. Meanwhile, a binary with no rule at all is still denied even during an audit run: Python's POST to the same host, with no Python entry in that rule, got permission denied and a DENIED line. Audit relaxes enforcement for the rule it is set on, not the deny-by-default around it.

The log itself is the product here. It is written in OCSF, the Open Cybersecurity Schema Framework, which is the event format security tooling already speaks, and every line names the binary, the destination, the rule that decided, and the engine that decided it, opa for the connection decision and l7 for the request decision. Useful flags: openshell logs demo --source sandbox --since 10m to see only the sandbox's own events, --level warn to cut the chatter, and --tail to stream. There is a setting, ocsf_json_enabled, meant to switch the events to JSON for a SIEM. I flipped it on a running sandbox and, in the minute I watched, the log view I was reading did not change; I did not chase it further. If you need JSON, set it before you create the sandbox and confirm where your driver writes it.

The advisor, the approval modes, and the trap I walked into

The advisor is the feature that answers Jake's question about sitting at an inbox, and it is also where the surprise on this page lives. Two things propose rules. The mechanistic mapper runs inside every sandbox whether you turned anything on or not: each denied connection becomes a draft rule for that host, port and binary, deduplicated on that triple, with a hit counter that climbs when the agent retries. Separately, if you enable agent_policy_proposals_enabled, the agent itself can ask for rules through a small local API, policy.local, using a guide OpenShell drops at /etc/openshell/skills/policy_advisor.md: it can read the current policy, list its recent denials, submit a proposal, and wait for a decision. The gateway is the single referee for both sources.

Every proposal, from either source, goes through the prover's proposal risk check before a human sees it, and then lands in the inbox with a confidence figure. Approval has two modes. manual, the default, puts everything in the inbox regardless of what the prover said. auto approves a proposal on its own only when the prover reports no new findings and there are no security notes attached; anything flagged still waits for you. You set it globally or per sandbox:

openshell settings set --global --key proposal_approval_mode --value auto --yes
openshell sandbox create --name auto --approval-mode auto --from my-image

What auto-approval actually approved

I created a sandbox in auto mode and ran one curl to the GitHub API. Denied, as always. Ten seconds later the mapper flushed its analysis, the proposal appeared, the prover said "no new findings", and the rule was approved without me, exactly as designed. After about twenty seconds the new version loaded and the retry succeeded. Then I sent a POST through the same rule, the kind of request a read-only rule refuses with a 403 from the proxy. It came back 401 from GitHub. The request left the sandbox, crossed the internet, and was refused by GitHub for lacking a token, not by OpenShell for lacking permission.

The reason is in the proposal itself, in a line that is easy to skim past:

Rule: allow_api_github_com_443
Binary: /usr/bin/curl
Confidence: 65%
Prover: prover: no new findings
Endpoints: api.github.com:443 [L4]

[L4]. The mechanistic proposal for a denied curl connection was a connection-level rule: any method, any path, no request inspection. It is the narrowest rule that would have let the denied connection succeed, and that is precisely all the mapper knows about a connection that never got far enough to send an HTTP request. The prover was also right: this rule involves no credential, so none of its four finding categories fire. Nothing malfunctioned. But the outcome is that auto mode, fed by connection denials, tends to grant more than a human writing the same rule would. In the same session, one of the four proposals I received, the one for Python, came back as an inspected read-only REST rule instead; I could not find a documented reason for the difference, so I will not invent one. The practical rule is simple: read the Endpoints line of every proposal, and treat [L4] as "allow everything on this host for this program." If that is more than the task needs, reject it with a reason and write the L7 rule yourself. And keep auto mode for sandboxes where no credential is attached and no host is sensitive.

Three smaller advisor facts that will save you a confused minute each. A new proposal for the same host, port and binary as a pending one supersedes it, and the older one is auto-rejected, so the inbox does not pile up with duplicates. Rejecting takes a --reason, and openshell rule history prints the proposed, approved and rejected timeline, so the next person knows why that door stayed shut. And rule approve-all and rule clear exist for the day you trust the inbox or want to empty it.

The prover: proving a policy stays inside a boundary

The prover is a separate binary, openshell-prover, that the Debian package installs next to the gateway, and it does something I have not seen in another sandbox tool. You give it two policies: a candidate, the one you want to run, and a boundary, the widest thing your organization permits. It uses an SMT solver to decide whether the candidate allows anything the boundary forbids, across filesystem access, process identity, Landlock settings, network reachability by binary, and REST method and path rules. It answers yes or no, and if no, it shows you one concrete example of what leaked.

I wrote a boundary that allows only read-only REST to the GitHub API for curl and Python, then checked three candidates against it:

$ openshell-prover check boundary.yaml --boundary boundary.yaml
result: within_boundary
coverage: domains=filesystem,network_l4,network_rest,process,landlock

$ openshell-prover check policy-advanced.yaml --boundary boundary.yaml
result: unsupported
reason: candidate policy rule 'pypi_audit' uses REST without enforced inspection

$ openshell-prover check policy-advanced-enforce.yaml --boundary boundary.yaml
result: exceeds_boundary
counterexample: network binary=- host=pypi.org:443 protocol=rest method=GET path=/

Three results, three lessons. A policy checked against itself is within boundary, which is the sanity test. My real policy, with the audit-mode pypi rule, was unsupported, because the prover models only enforced inspection; an audit rule is a promise the sandbox is not keeping, so the solver refuses to reason about it. And the same policy with pypi switched to enforce was exceeds boundary, with the counterexample spelled out: a GET to pypi.org, which the boundary never allowed. The exit codes are 0, 1, 2 and 3 for within, exceeds, error, and unsupported or inconclusive, and --output json prints a structured report with the counterexample as fields. That is a CI job: keep the boundary in the repository, run the prover on every policy change, fail the build on anything but zero.

Know its edges. GraphQL, MCP and JSON-RPC rules are not modeled, and neither are path comparisons that could be fooled by symlinks, so a policy using them returns unsupported. And "within boundary" means the candidate grants nothing beyond the boundary in the parts the prover checks. It does not mean the policy is narrow, or right for the task, or that the sandbox enforced it. It is a proof about the document, not about the world, and it is still worth having.

Inside the advisor, the prover runs a different check on every proposal, comparing what the sandbox can reach with and without the proposed rule, and it reports four kinds of finding: reach to link-local or cloud metadata addresses, uninspected traffic from non-HTTP tools such as ssh or nc to a host that carries a credential, a credential becoming usable at a new host and port, and a new HTTP method on a destination that already carries a credential. Any one of those blocks auto-approval. Notice what is not on the list: a plain connection to a public host with no credential attached. That is why my curl proposal passed.

Credentials: providers, placeholders and credential binding

I did not attach a paid provider for this guide, so this section is from the documentation and the tool's own help text rather than from a run, and it is short on purpose. The design, though, is the part of OpenShell I would most want a team to understand, because it is what makes a leaked agent survivable.

A provider holds a real credential, an API key or a token, in the gateway. A profile describes a service: which environment variable the credential lives in, which hosts, ports and paths it may be sent to, and which binaries are known to use it. At sandbox start, the agent's environment gets an opaque placeholder instead of the key. When a request carrying the placeholder reaches the proxy, and only if the destination matches the profile's declared endpoints, the proxy swaps in the real value on the way out. Send the placeholder anywhere else, even a host your network policy allows, and the proxy answers 403 with the reason credential_endpoint_mismatch. A network rule never widens where a credential can go; the profile does, and only the profile. So an agent that dumps its own environment, or is tricked into posting it somewhere, hands over a token that is worthless outside the sandbox.

In YAML, a profile with no endpoints of its own is bound to a concrete destination with credential_binding: {provider: name} on the endpoint, and there are flags for the awkward cases: request_body_credential_rewrite to replace placeholders inside JSON bodies, a WebSocket equivalent, and credential_signing: sigv4 with a signing_service such as bedrock or s3 for AWS APIs, which need the request signed rather than a header set. That last one is how you would build a sandbox for agentic AI on AWS Bedrock, letting the agent call Bedrock with credentials it never sees; if you have fought Bedrock access errors before, the idea of the gateway holding the AWS keys will appeal to you. The gateway can also refresh tokens itself, through OAuth refresh, Google service-account JWTs or AWS STS, keeping the refresh material gateway-only and injecting only the short-lived access token. The commands are openshell profile import, openshell provider create --name x --type y --credential KEY, and --provider x on sandbox create; the beginner guide walks the OpenRouter version.

OpenShell MCP rules, plus GraphQL, JSON-RPC and WebSocket

Agents increasingly talk to tools over MCP, the Model Context Protocol; an MCP server is a tool server the agent calls over HTTP with JSON-RPC bodies, and OpenShell can write rules against those calls. I did not have an MCP server to point a sandbox at, so what follows is the documented schema, not a run; the shape is the same as the REST rules you saw work above, and these all require policy set with full YAML. An MCP endpoint looks like this:

  tools_server:
    endpoints:
      - host: mcp.example.com
        port: 443
        protocol: mcp
        enforcement: enforce
        rules:
          - allow:
              method: initialize
          - allow:
              method: notifications/initialized
          - allow:
              method: tools/call
              tool:
                any: [search_web, list_issues]
        deny_rules:
          - method: tools/call
            tool: send_email
    binaries:
      - path: /usr/local/bin/my-agent

OpenShell parses the JSON-RPC envelope in the request body, validates the known MCP request and notification shapes, and matches on the method and, for tools/call, the tool name, with globs allowed in the tools/ family and a strict tool-name pattern on by default. It also refuses any request to an MCP endpoint that carries an Upgrade header, in every enforcement mode, because a raw relay would bypass per-request rules. Be clear-eyed about the limits, which the documentation states plainly: tool arguments are not matched, one denied call in a batch fails the whole batch, and enforcement is one-directional, on the request bodies the sandbox sends. Responses, including server-sent event streams, are relayed but not parsed. So OpenShell can stop an agent calling send_email; it cannot stop it calling search_web with a query that contains your customer list. That is a job for the tool server or for a middleware, and there is a small industry forming around it.

GraphQL rules match operation type, operation name and top-level fields, with persisted queries denied unless you register them, and a 64 KiB body limit by default. JSON-RPC rules match the exact method name only. WebSocket rules cover the upgrade request and client text frames; binary frames and everything the server sends are not inspected. In every case the same precedence holds: a matching deny beats any allow.

The DNS exfiltration hole, and how it was closed

Every sandbox that controls the network eventually meets the same attack, and OpenShell met it in public. On May 5, 2026, a researcher opened issue #1169 on the NVIDIA OpenShell GitHub repository. The finding: the proxy checked TCP and HTTP against policy, but name resolution went out over UDP port 53 on its own, so a program could encode data into subdomain labels and have it delivered to an attacker's nameserver as a lookup, before the TCP connection was ever refused. The report showed a file leaving a sandbox that way with tools the base image already had. A maintainer replied on May 12 that a fix was up, and pull request #1329 was merged on May 15: it removed the last place the sandbox itself resolved a denied hostname, a helper in the proposal mapper that had looked up names to decide whether they pointed at private addresses. Proposals no longer carry auto-filled allowed_ips, and permitting a private range is now a deliberate two-step: the SSRF guard denies it first, then an operator adds the range by hand.

I did not reproduce any of that, and this page will not show you how. What I can show is what name resolution looks like from inside a 0.1.2 sandbox today, because it is visible in ordinary logs. When I looked up example.com from a sandbox with no rule for it, the answer was 198.18.0.2, an address from a reserved benchmarking range, and the log said:

NET:REFUSE [MED] DENIED example.com [reason:policy_dns_ineligible]
CONFIG:OBSERVATION [INFO] Policy DNS staged unapproved name example.com synthetic=198.18.0.2 for TCP policy review

The sandbox never asked a real resolver. The supervisor handed back a synthetic address that only means "I have noted you want this name," carried the name forward to the TCP decision, and refused the connection there. Only after a rule allowed api.github.com did a real lookup happen, on the outside, and the log recorded it: "Policy DNS mapped api.github.com resolved=20.207.73.85 synthetic=198.18.0.4." The mapping is short-lived and becomes stale on any policy change. A program inside cannot learn where a server lives until you permit the server, and a name that is not in policy never leaves the machine as a query.

Two cautions remain, and the documentation is candid about both. Wildcard hosts authorize lookups for every name that matches, and a label-encoding channel needs exactly that, so a rule like *.internal.example reopens a narrow version of the problem for that domain; keep wildcards to hosts you control, and prefer exact names. And a rule that allows a host to be reached by name allows the private address it resolves to, which is why wildcards and hostless rules must carry allowed_ips, and why the proxy hard-blocks loopback, link-local, the cloud metadata address and the Kubernetes control ports regardless of policy. If you take one thing from this section: read the issue and the pull request yourself; they are short, and they are the most honest documentation of the tool's threat model that exists.

Filesystem, process and resource limits

The network gets the attention, but a sandbox for an agent that edits files is mostly about the filesystem. The default is what you saw in the YAML: seven read-only paths, /tmp and /dev/null writable, and the working directory writable. Paths must be absolute, no .., at most 256 of them, and / itself can never be writable. Landlock enforces it, and Landlock is why the walls are fixed at start: the rules are applied to the process tree before the agent runs, and the kernel does not let a sandboxed process widen its own rules afterward.

Three behaviors I checked because they decide how you actually work with an agent over days:

  • Stop and start keep everything. I wrote a file in /sandbox and one in /tmp, stopped the sandbox, started it, and both were there, with the policy version unchanged. Stopping is for pausing, not resetting.
  • Upload and download are the door for code. openshell sandbox upload demo ./project lands in the working directory as /sandbox/project, with .gitignore respected unless you pass --no-git-ignore. openshell sandbox download demo /sandbox/keep.txt . brings a result home. The agent's edits stay inside until you choose to fetch them, which is the right default for a week-long job on a client repo.
  • The user is yours to set, once. With a complete policy including process: {run_as_user: "1500", run_as_group: "1500"}, the agent ran as uid 1500 and could write its workspace. Root is refused outright, at create time. The default is the image's USER or 1000.

Resource caps work the way container caps work, and they are real. openshell sandbox create --cpu 1 --memory 512Mi produced a container with exactly one CPU's worth of quota and a 512 MiB memory limit in the container's host configuration, a process limit of 2048, and no network interface at all, net=none, which is the physical version of deny-by-default: the only way out is the supervisor's proxy. --gpu requests a device through NVIDIA's container device interface for the day you sandbox a model that needs one. For a team, openshell sandbox template create bakes image, CPU, memory, GPU, labels and environment into a named template so every sandbox starts from the same shape.

Global policy: locking a fleet, and the two ways it bit me

A gateway administrator can set one policy for every sandbox on the gateway. While it is in force, sandbox-level changes are refused, proposal approvals are refused, and provider-contributed rules are suppressed, so nobody on the team can quietly open a door the organization closed. It is the single most important feature for anyone rolling this out beyond one laptop, and it comes with two sharp corners I found the hard way.

openshell policy set --global --policy global.yaml --yes
✓ Global policy configured (hash: 909aadcdcd07, settings revision: 1)

openshell policy update demo --add-endpoint example.org:443:read-only:rest:enforce
Error: policy is managed globally; delete the global policy before using `openshell policy update`

That part behaved perfectly. The first corner: removing the lock needs --yes when you are not at an interactive terminal. My script ran openshell policy delete --global without it, the command printed "global setting deletes require confirmation; pass --yes in non-interactive mode", my script did not notice, and the lock stayed on. The second corner followed from the first, and it is the one to remember. With that global policy in place, which contained only a network rule and no filesystem section, every new sandbox I created went to phase Error with the condition "Waiting for effective configuration validation before workload activation," and a sandbox I stopped and started did the same. After I deleted the global policy properly, the new sandboxes came up fine, but the one that had broken under the lock stayed broken and had to be deleted and recreated. So: write a global policy as a complete file, starting from the full default; delete it with --yes in scripts and check the exit code; and expect sandboxes that started under a bad global policy to need recreating, not restarting.

Runtimes: container, microVM, Kubernetes, and Windows

Everything on this page ran in a rootless Podman container, which is the default on Linux and the right choice for a laptop. It is not the only boundary the gateway can use, and the choice is a one-line setting, compute_driver in the gateway's TOML file or the environment variable OPENSHELL_COMPUTE_DRIVER. I did not run the other drivers; this table is from the documentation.

Driver Boundary Needs Where
podmanRootless containerPodman 5, cgroups v2, user socketLinux, WSL 2
dockerContainerDocker daemonLinux, macOS
vmOne lightweight VM per sandboxKVM on Linux, Apple Hypervisor on macOS; opt-in, never auto-detectedLinux, macOS
kubernetesPod in a sandbox namespaceAgent Sandbox controller, a CNI that enforces NetworkPolicyAny cluster
mxcNative Windows, via Microsoft execution containersListed as coming soonWindows 11

When would you leave containers? When the agent will run code you truly do not trust, such as evaluating third-party MCP servers or running a model that writes and executes its own scripts; a kernel shared with the host is a larger surface than a VM boundary, and the vm driver exists for that. Two practical notes from the project's issue tracker: the macOS VM driver shipped in a release tarball without the Apple hypervisor entitlement and fails to create a VM unless installed through Homebrew, and native Windows support without WSL is a proposal, tracked as issue #2050, not a thing you can run today. On Windows, WSL 2 with Podman inside remains the path, and it behaves like my Kali laptop did.

The rough edges, including the one I got wrong last time

Collected in one place, because the difference between a good evening and a bad one with a version 0.1 tool is usually one sentence nobody wrote down.

  1. Approve returns before the reload. rule approve comes back in a second; the sandbox took about twenty seconds to load the new version. Retry too early and you get the old denial, and you may conclude, as I did in the beginner guide, that something is broken. Watch policy list for Loaded.
  2. A policy file replaces the default. Omit filesystem_policy and the sandbox exits with code 1 at start. Start every custom policy from the full default.
  3. Filesystem and process changes need a recreate. policy set will accept them and policy get will show them, and nothing changes until the sandbox is recreated.
  4. Auto-approved rules from connection denials are L4. Any method, any path. Read the Endpoints line; write the L7 rule yourself if it matters.
  5. Audit-mode violations are logged as ALLOWED. Review by rule name, not by grepping for DENIED.
  6. The enforcement default is audit. A rule you forgot to mark enforce is logging, not blocking. Say enforce explicitly, every time.
  7. Global policy delete needs --yes in scripts, and sandboxes that broke under a bad global policy need recreating.
  8. Your image must not run as root, and if you set a custom uid the image does not know, whoami will grumble; the sandbox is fine.
  9. Denials look different depending on the path. My curl got (7) Couldn't connect because the sandbox used transparent interception; the docs' example shows (56) 403 from proxy after CONNECT from a sandbox using the proxy environment variables. Both mean policy denied; the log line is the same.
  10. The prover will not reason about audit rules, GraphQL, MCP or JSON-RPC. Prove the enforced REST and connection layer, and review the rest by eye.

None of these is a reason not to use the tool. They are the texture of something young and moving fast, and every one of them was discoverable in an evening. I would rather you meet them here than at eleven at night with a client's repository in the sandbox.

For teams and IT: a rollout you can defend

Ethan's real question was never "does it work," it was "can I put my name on it." Securing AI agents with zero trust is the phrase the vendors use; deny-by-default with per-program rules is what it means in practice. Here is the shape of an answer, built from the pieces above, in the order I would do it.

  1. Write the boundary first. One YAML file, kept in version control, describing the widest access any agent at the company may have. It is the document your security review actually approves.
  2. Prove every policy against it in CI. openshell-prover check candidate.yaml --boundary boundary.yaml --output json, fail on any exit code but zero. A policy that exceeds the boundary never reaches a gateway.
  3. Roll out in audit, read the log, then enforce. Give the agent its real work for a few days with rules in audit mode. Build the enforced policy from what it actually reached, not from what you imagined.
  4. Lock the fleet with a global policy, complete file, and keep approval mode manual on any sandbox that carries a credential.
  5. Ship the log to where your findings already live. The OCSF events belong in the same place as your other security findings; if you already use something like AWS Security Hub or a SIEM, that is the destination. The policy version hash in each decision is your change record.

Security for AI agents, in other words, is mostly a policy problem with a small runtime underneath it. Be honest about what this does not cover, because your reviewers will ask. OpenShell governs one sandbox at a time. It does not give agents an identity that persists across sandboxes, it does not mediate agent-to-agent traffic, and it does not tell you which model answered a request; people building fleet tooling on top of it make exactly that argument, and they are right that it is by design. At the protocol layer, it sees MCP method and tool names but not arguments and not responses. And a shared kernel is still a shared kernel: for untrusted code, use the VM driver. It does not stop a prompt injection from happening, but it caps what a successful one can reach, which is the part you can actually control. What it does do, and does well in my testing, is turn "we let engineers run agents" from a blind spot into a stream of decisions you can read, prove, version and hand to someone else. That is more than most teams have today.

What is an OpenShell policy?

A YAML file with version 1 and five sections: filesystem_policy, landlock, process, network_policies and network_middlewares. The first three are fixed when the sandbox starts and enforced by Landlock and the container engine. The network sections reload live and are enforced by the OpenShell policy engine: OPA for connection decisions and an L7 inspector for request decisions, both inside the supervisor's proxy. Anything the policy does not allow is denied.

What is the difference between openshell policy set and openshell policy update?

Update merges small changes into the live network policy: add an endpoint, add an allow or deny line, remove a rule. Set replaces the whole policy with a file, and is required for middleware, GraphQL, MCP and JSON-RPC rules, TLS skipping, credential bindings and request signing. Both create a new numbered, hashed policy version, and both accept --wait to block until the sandbox has loaded it.

How do OpenShell network rules combine when several match?

They are not an ordered list. Every matching allow rule adds the access it grants, and a single matching deny rule wins over all of them wherever it appears. Two rules for the same host with different binaries therefore both apply. If a retry still fails after adding a rule, check that the new policy version shows Loaded first; the reload takes several seconds.

What does the OpenShell policy prover do?

It checks whether a candidate policy allows anything a boundary policy forbids, across filesystem, process, Landlock, network reachability by binary, and REST method and path rules, using an SMT solver. It returns within_boundary, exceeds_boundary with a counterexample, error, or unsupported, with exit codes 0 to 3, and has a JSON output for CI. Inside the advisor it also risk-checks every proposed rule.

Is OpenShell auto-approval safe to turn on?

It only approves proposals the prover finds clean and that carry no security notes, so credential-related rules still wait for you. But a proposal drafted from a denied connection is a connection-level rule allowing any method and path to that host for that binary, which is broader than a hand-written read-only rule. In my test a POST went straight through such a rule. Keep auto mode for sandboxes with no credentials and no sensitive hosts, and read the Endpoints line of every proposal.

What is the difference between audit and enforce in OpenShell?

Enforce blocks a request that violates the rule and returns a policy_denied error to the agent. Audit lets the request through and logs it, so you can see what an agent really needs before you tighten the rule. Audit is the default when you leave enforcement out, so state enforce explicitly. Audit relaxes only the rule it is set on; a binary or host with no rule is still denied.

Does OpenShell inspect HTTPS traffic?

Yes. The supervisor terminates TLS with a per-sandbox ephemeral certificate authority, inspects the request against the rules, and reopens TLS to the real server. Inside the sandbox the usual variables such as SSL_CERT_FILE, REQUESTS_CA_BUNDLE and NODE_EXTRA_CA_CERTS point common tools at that authority. An endpoint can set tls: skip to disable inspection, at the cost of request-level rules.

Can OpenShell control which MCP tools an agent calls?

Yes, with protocol: mcp on the endpoint and rules matching the MCP method and, for tools/call, the tool name, including deny rules for specific tools. It parses the JSON-RPC body the sandbox sends. It does not match tool arguments, does not parse responses or server-sent event streams, and refuses any upgrade request on an MCP endpoint. Those rules need policy set with full YAML.

How does OpenShell handle DNS and prevent DNS exfiltration?

A sandbox never queries a real resolver for a name that policy does not allow. It receives a synthetic address from a reserved range, the name is carried to the connection decision, and the connection is refused and logged. Real resolution happens outside only after a rule allows the host. The DNS exfiltration issue reported in May 2026, issue 1169, was fixed by pull request 1329 that month. Wildcard hosts still authorize lookups for every matching name, so prefer exact hostnames.

How do I run an AI agent as a specific user in OpenShell?

Add a process section with run_as_user and run_as_group as quoted numbers to a complete policy file and pass it with --policy at create time. Root is refused. The policy must also contain the filesystem_policy block, because a policy file replaces the default rather than extending it; without it the sandbox exits at start.

Can I limit CPU and memory for an OpenShell sandbox?

Yes. openshell sandbox create --cpu 1 --memory 512Mi sets a CPU quota and a memory limit on the container, and --gpu requests a GPU through the NVIDIA container device interface. Templates created with openshell sandbox template create bake these into a reusable shape for a team.

What is a global policy in OpenShell?

One policy a gateway administrator applies to every sandbox with openshell policy set --global. While it is active, sandbox-level policy changes and proposal approvals are refused and provider-contributed rules are suppressed. Write it as a complete file; in scripts, remove it with openshell policy delete --global --yes and check the exit code.

How to secure AI agents that run commands on your machine?

Run them in a sandbox that denies everything by default and grant access one narrow rule at a time: a read-only filesystem outside the workspace, no network except named hosts for named programs, credentials held outside the sandbox and injected only for their own service, and a log of every decision. That is the model OpenShell implements; the important habit is audit first, enforce second, and read the log.

Can AI agents be trusted with a terminal?

Not on their own, and OpenShell is built so you do not have to. The agent runs where it cannot reach anything you did not name, every request is judged by a rule it cannot edit, and every decision is logged. Trust moves from the model to the walls, which is a far better place to keep it. How to secure agentic AI, in one line: assume it will try everything, and make almost everything impossible by default.

What is a sandbox in cybersecurity?

A sandbox in cyber security is an isolated environment where untrusted code can run without reaching the real system, network or data around it. Classic sandboxes detonate suspicious files; an AI agent sandbox does the same for a program that is supposed to be helpful but cannot be fully trusted, and adds fine-grained rules so the agent can still do useful work through approved doors. The sandbox meaning in cybersecurity has not changed; the thing inside it has.

What does sandbox env in CLI for AI agents mean?

An agent sandbox you create and control entirely from the command line, rather than a hosted sandbox-as-a-service. OpenShell is that: openshell sandbox create makes the cell, openshell policy writes the rules, openshell rule handles proposals, and openshell logs reads the audit trail, all from a terminal on your own machine.

Can an AI agent escape an OpenShell sandbox?

An AI agent sandbox escape needs an open door, and the ordinary ones are shut. It runs as a non-root user with empty kernel capabilities and a seccomp filter, on a Landlock-restricted filesystem, in a container with no network interface, behind a proxy that resolves names only for allowed hosts. The copied-binary and symlink tricks are handled by real-path identity. No sandbox is a proof, which is why the VM driver exists for genuinely untrusted code and why every attempt is logged.

What is an MCP server, and why sandbox it?

An MCP server exposes tools, such as search, file access or email, that an agent can call over the Model Context Protocol. Because those tools act in the real world, a sandbox should decide which of them the agent may call. OpenShell does that by name with protocol: mcp rules; argument-level control and response inspection are outside its scope today.

Does the OpenShell prover work with MCP or GraphQL rules?

No. The prover models enforced REST rules and connection-level network rules, plus filesystem, process and Landlock settings. A policy containing GraphQL, MCP or JSON-RPC rules, or REST rules in audit mode, returns unsupported. Prove the parts it covers and review the rest manually.

πŸ“š ALSO READ

Running agents for real means credentials, cloud APIs and a model behind them. These connect the dots:

📌 Bookmark this if you are writing an agent policy for a team this quarter.

If you got this far, you now know more about how this tool decides than most of the people who will be asked to approve it. Use that gently. The point of every section above is not that agents are dangerous; it is that you can make them boring, in the best sense, by deciding in advance what they may touch and keeping the receipts. Start with the boundary, prove against it, audit before you enforce, and read the Endpoints line before you say yes. Jake still does not want to sit at an inbox all day, and with a good boundary and a prover in CI, he will not have to. If something on this page does not match what you see on your screen, write in; this tool is at version 0.1 and I would rather correct a paragraph than let it mislead you, as the beginner guide shows.

📌 If you keep one line from this page

Approve the rule you read, not the rule you assumed; an L4 line is a whole host.

Write the boundary, prove against it, audit first, enforce second, and treat every proposal's Endpoints line as the contract.

Revision note. Written September 29, 2026, from runs that day on Kali Linux, kernel 7.1.5, Podman 5.8.6, OpenShell 0.1.2 and openshell-prover 0.1.2. The policy YAML, path and deny rules, binary identity tests, audit-mode run, auto-approval run, prover checks, process and resource limits, stop and start behavior, and the global policy lock are all first-hand, with log lines pasted as printed. The credentials, MCP, GraphQL, JSON-RPC, WebSocket and runtime-driver sections are from the project's documentation and tracker and are marked as such in the text. The DNS section is from the public issue and pull request; nothing in it was reproduced. This page also corrects a claim in our beginner guide about overlapping rules, which was a timing error on my side.

Related