Approve outcomes, not commands. Start a long agent task without permission
prompts and do something else. The agent works in a copy-on-write branch of your
workspace and $HOME, git push and the commands you name wait in an outbox, and
every host it reaches is logged. When you come back, one review shows what changed and what is waiting:
apply it, take it onto a git branch, or throw it away.
Install · What you get · How it compares · Threat model · FAQ · Docs
Note
Early v0, Linux first; macOS is a prototype. By default airbag is not a VM: the native backend shares the host's kernel, and whatever the agent reads is still sent to the model API. It guards against accidents and casual exfiltration by an agent you let run without prompts; for code that may try to break out, use a VM (airbag's experimental microVM backend has its own limits). See the threat model.
$ curl --proto '=https' --tlsv1.2 -fsSL https://github.com/getjump/airbag/releases/latest/download/install.sh | sh
$ airbag doctor # can this machine run airbag?
$ cd your-project
$ airbag run -- claude --dangerously-skip-permissions
$ airbag run -- codex --dangerously-bypass-approvals-and-sandbox # or Codex
$ airbag review
$ airbag apply # or: apply -i, apply --branch NAME, or: airbag discard
$ airbag rollback # undo the last applyRequirements, go install and Nix: Install.
The demo runs the real Claude Code; the "model" is test/mockapi playing a fixed
script, so it is repeatable without an account (demo/demo.sh). agent ▶ lines are
the calls the model makes, agent ◀ what the agent sends back. The same as a GIF:
demo/demo.gif. More scenes, one GIF each, in demo/: sandbox, codex, ask,
apply, mirror (demo/scenes.sh NAME; demo/render.sh NAME makes the SVG).
- A branch of the world. The workspace and
$HOMEare copy-on-write branches, the rest of the host is read-only. Nothing the agent writes to the workspace reaches your files beforeairbag apply, and almost nothing it writes to$HOME. - One way out. Traffic leaves only through airbag's proxy: model APIs and the hosts you allow, every host logged. Packages come through a read-only mirror.
- An outbox.
git pushand the commandsdefer:names wait there and run on the host after review. - No credentials.
~/.ssh,~/.aws,ghand credential-like variables are hidden; a token you bind to hosts reaches the agent as a placeholder. - Watched secrets. The first read of
.envor a key labels the session and narrows egress: only model APIs stay reachable. - Less kernel to attack. A seccomp filter
refuses io_uring,
bpf, the keyring and more;--strictforbids new user namespaces. - One review, then your call. Apply all of it, part or none,
or onto a new git branch;
airbag rollbackundoes an apply. - Rules over effects. CEL rules on observed
effects (
net.connect) and predicted ones (net.egress,fs.delete) answerallow,denyorask. - More than one run.
airbag run --session last -- claude --continueruns the agent again on the same branch.
Isolation is not the difference. Claude Code's and Codex's built-in sandboxes run on bubblewrap on Linux and Seatbelt on macOS, and airbag uses the same kernel features. What differs is when you decide, and what you get to see.
Say you ask an agent to clean up a repository, and it deletes src, reads
.env, tries to post it to a paste site, appends a line to ~/.bashrc and
pushes.
- Built-in sandbox: the agent stops for approval as it goes: a write outside the workspace, a new domain. You answer prompts mid-run and never see the whole result at once, and writes inside the workspace land in place.
- airbag: the agent runs without stopping in a branch of your machine.
Afterwards one review shows all of it:
srcdeleted,.envread, the upload blocked, the~/.bashrcline flagged as persistence, the push waiting in the outbox.airbag discard: the files and~are as they were, and the push never left. (A request to an allowed host would have happened when it was made; see below.)
| Built-in sandbox (Claude Code, Codex) | Dev container | VM or microVM | airbag | |
|---|---|---|---|---|
| Isolation | bubblewrap on Linux, Seatbelt on macOS | a container; shared kernel | its own kernel | namespaces, overlayfs and seccomp around the whole agent; shared kernel by default (experimental gVisor and microVM backends on Linux) |
| You decide | during the run, at each prompt | before: mounts and network | before: what goes in | after: one review of the whole run |
| Writes to the workspace | land in place | land in place (bind mount) | stay in the VM until you copy or merge them out | stay in a branch until apply; rollback undoes an apply |
git push |
runs if the network allows it | runs if the network allows it | runs if the network allows it | waits in the outbox until review |
| Setup | built in | Docker and a config | a VM image and its tools | one binary, Linux 5.12+ |
As of October 2026, from each project's documentation. Corrections welcome: open an issue. The full comparison with the tools that also let you decide after the run (nono, try, AgentFS, Docker Sandboxes, Claude Code checkpoints), and why airbag does not run on bubblewrap: docs/comparison.md.
| Until you decide | Examples | |
|---|---|---|
Stays local until airbag apply |
the workspace and $HOME |
edits, deletions, new files, a line in ~/.bashrc |
| Waits in the outbox, runs after review | what airbag intercepts | git push, and the calls defer: names, such as gh pr create or npm publish |
| Decided when it happens, by the allowlist and policy | every other network request | model API calls, a request to a host you allowed, packages through the mirror; ask holds a call until you approve it |
A call to an outside service cannot be held or undone after the fact, so for those the allowlist and policy decide beforehand. Review shows that they happened.
airbag protects against accidents and casual exfiltration by an agent you let run
without prompts. By default it is not a VM: the native backend shares the host's
kernel (the experimental --backend=microvm boots a guest kernel instead, with the
limits in docs/runtime-options.md), and whatever the
agent reads is still sent to the model API. A bound credential keeps its value from the agent,
not its use: through the bound hosts the agent can do what the token allows.
- The kernel. With the native backend, the seccomp filter makes the shared
kernel a smaller target, not a VM boundary.
airbag doctorreports the host sysctls that harden the rest. - The network. For hosts without a credential the proxy decides from the name
the client asks for and does not see inside TLS, so a broad allowlist entry
(
github.com) is a way for data to leave. Allow narrow names, and where that matters put anaskrule onnet.connectfor the broad ones. - The host side. The proxy, the forwards and the control socket run in airbag's process on the host. Connections that stop carrying data are closed and their number is capped; bandwidth and the rate of new connections are not limited.
The whole threat model, with the limits in numbers: docs/threat-model.md.
It is if you let Claude Code or Codex run long tasks without permission prompts and
want to see the whole result before it reaches your files, ~ or a remote, on
Linux, with your own toolchain and, by default, without a VM. It is not for you if:
- you need a hard boundary against code that tries to break out: use a VM or a microVM (airbag's own experimental gVisor and microVM backends have the limits in docs/runtime-options.md);
- what the agent reads must not reach the model provider: airbag does not change what is sent to the model;
- you are on Windows, or need macOS today: the macOS port is a prototype, and docs/macos.md has the Linux VM setup that works;
- you need a hosted runtime for many users: airbag runs on your machine.
Is airbag a VM, or a hard security boundary?
Not by default. The native backend puts Linux namespaces, overlayfs and a seccomp
filter around the agent, on the host's kernel. That guards against accidents and
casual exfiltration; it is not built to hold code that tries to break out. For that,
use a VM or a microVM. On Linux, the experimental --backend=gvisor and
--backend=microvm run the agent on gVisor's kernel or in a Firecracker VM, with
the same review, apply and outbox, but a narrower profile and their own limits:
docs/runtime-options.md.
How is it different from the sandbox built into Claude Code or Codex?
Not in isolation: airbag uses the same kernel features they do. A built-in sandbox asks as the agent goes, and writes inside the workspace land in place. Under airbag the agent runs without stopping in a branch, and you decide afterwards, in one review, what to keep.
Does the model still see my files?
Yes. Whatever the agent reads is sent to the model API, as it is without airbag. Known secret values are masked in the output of shell commands, but the agent's own file tools are not filtered, so a secret the agent reads can reach its model. The label a secret read puts on the session stops it from going anywhere else.
What happens to a request to a host I allowed?
It happens when the agent makes it: a call to an outside service cannot be held or
undone after the fact, so the allowlist and rules decide beforehand. An ask rule
blocks the call, and its retry passes once you run airbag approve. Review lists
every host that was reached.
Which agents does it work with?
Any command: airbag run -- AGENT [ARGS...]. Claude Code and Codex also get
airbag's hooks, so the review shows which tool call changed which file. Gemini CLI,
Aider, OpenCode and others run in the sandbox without that attribution.
Does it run on macOS or Windows?
Linux first. On macOS there is a native prototype (Seatbelt around the agent, an APFS clone as the branch) and a Linux VM setup that works today; see docs/macos.md. There is no Windows build; under WSL2 the Linux build runs, and an informational CI job runs the end-to-end tests there.
What does it cost in speed and disk?
Not measured yet; docs/evaluation.md is the plan. A session
keeps what the agent wrote under /var/tmp/airbag-$UID (AIRBAG_HOME moves it)
until you discard it, and packages the mirror fetched are cached across sessions.
Can I undo an apply?
The last one: airbag rollback puts back what was there, and the agent's changes
go back into the session to apply again or discard. A file you edited after the
apply is left as it is. A push that already ran is not undone.
$ curl --proto '=https' --tlsv1.2 -fsSL https://github.com/getjump/airbag/releases/latest/download/install.sh | sh
$ airbag doctorThe script installs the latest release for Linux or macOS 13 and later (amd64,
arm64) into ~/.local/bin, without sudo. It checks the archive's SHA-256 against
the release's checksums.txt, and when cosign or a logged-in gh is installed,
that the release workflow built it. | AIRBAG_VERSION=v0.1.0 sh picks a release,
and | AIRBAG_VERIFY=require sh refuses to install without the second check.
docs/verify.md shows how to verify a release by hand or rebuild
it. Or from source, with Go 1.27.1 or newer:
go install github.com/getjump/airbag/cmd/airbag@latest.
With Nix: nix run github:getjump/airbag -- doctor, or add the flake's
packages.<system>.airbag to your configuration. nix flake check runs the unit
tests, nix develop gives a shell with Go and the test tools.
One static binary, no daemon, no Docker. Needs Linux 5.12+ with unprivileged user
namespaces. On Ubuntu 23.10+ AppArmor restricts them; airbag doctor prints the
one-time profile to install. On macOS there is a native prototype (Seatbelt
around the agent, an APFS clone as the branch), tested in CI on macOS 15 and 26 but
not yet used on real work, and the Linux VM setup that works today; see
docs/macos.md.
Early v0, not yet tried by anyone outside the project. Working on Linux: sandbox,
branch, proxy with allowlist and address checks, mirror, outbox for git push and
defer: commands, review (text, --attention, --json), diff, apply with conflict
check (all or nothing), apply --branch, rollback, run --session, tcp://
forwards, credentials through placeholders, discard. On macOS: a prototype (see above).
How the agents' hooks and the shell models work, the tests, the effect log and what
is not done yet: docs/status.md.
Deliberately left for later, each with the reason and what would bring it back
(docs/roadmap.md): placeholders in .env files, TLS termination
for model APIs, data flow labels per value, syscall-level control, savepoints per
tool call.
Before more features comes a measurement on real work against the alternatives:
docs/evaluation.md.
- What the agent gets: files, network, outbox, secrets, kernel, sessions
- Review and apply:
--attention,--json, rollback,apply --branch - Policies: rules over effects, deferred commands, bound credentials
- Runtime policy: opt-in file and exec checks on Linux, their audit and limits
- Threat model: what airbag protects against, and what not
- Execution boundaries:
airbag capabilities,--require-isolation, no fallback; optional gVisor and microVM runtimes on Linux - How airbag compares: nono, try, AgentFS, Docker Sandboxes, agentsh
- Status in detail: hooks, shell models, tests, the effect log
- Roadmap and decisions: what is deferred on purpose, and why
- Evaluation plan: how to tell whether airbag is worth using
- Embedding airbag's components: the operation contracts, policy evaluator, outbox and credential proxy as public Go packages, the same code the CLI runs
- macOS, bubblewrap as a backend, manual test with a real login, described operations
CONTRIBUTING.md says how to build and test airbag, and where help changes the most. Runs on real work, and issues about review noise, false flags and tools that break in the sandbox, are the most useful input now.
Please report security problems privately through a GitHub security advisory, not a public issue; see SECURITY.md.