Open-source, model-agnostic, security-first autonomous coding agent.
π Documentation β https://riz007.github.io/larb/
New here? Start with Getting started and the Security model before your first run β Larb runs autonomously and executes commands on your machine.
Larb is a terminal-native autonomous engineer you can point at any model, extend with governed skills, and run without rate-limit cliffs or vendor lock-in β with a trust model that makes opening an untrusted repo safe by default.
β οΈ Status:0.1.0-alphaβ early, BYO-key developers only. APIs, config, and the skill format may change without notice. Review a run's actions rather than leaving it fully unattended, and read Known limitations before relying on it. Feedback and issues very welcome.
- Trust-before-anything boot β no executable config is read and no network call is made before you make a trust decision for a directory.
- Capability-gated tools with a fine-grained permission engine (allow once / for the session / always / deny).
- Append-only audit log of every model call, tool call, and grant.
- Hard spend governor β live token/$ accounting with limits that halt the agent rather than just warn.
- Model-agnostic provider abstraction β point Larb at Anthropic, OpenAI, DeepSeek, Gemini, Groq, Mistral, xAI (Grok), OpenRouter, Together, Perplexity, or a local Ollama model by changing one line of config. See Providers below.
- Agent orchestration loop β plan β act β observe β verify, with a mandatory verification step (lint/build/test) before a task is "done".
- Real sandboxed execution β commands run in a rootless container (docker/podman) with the project bind-mounted, host secrets withheld, and networking off by default; falls back to a reduced-isolation host subprocess when no runtime is present, and tells you which is active.
- Governed network egress β the agent's only in-process network path is an
http_fetchtool gated per-host by thenetcapability (default-deny, every host approved and audited); in containerallowlistmode, command egress is routed through a host proxy that permits only allow-listed hosts. - Durable run state β every run is snapshotted;
larb runslists them andlarb resumecontinues an interrupted one exactly where it stopped. - Incremental repo map + inspectable markdown memory.
- Governed skills, installable from a directory, an https tarball, or a git URL β signed/manifested, with install β trust (unsigned β tightest sandbox).
- MCP (Model Context Protocol) servers β plug in external tools (filesystem,
GitHub, custom company servers) via
[[mcp]]in your global config. Each remote tool is a permission-gated, audited tool (mcp__<server>__<tool>); servers connect only inside a run, after trust. Inspect them withlarb mcp/larb mcp probe. AGENTS.mdproject instructions β Larb readsAGENTS.md(and.larb/AGENTS.md) and injects it as advisory guidance, so a repo can describe its build/test commands and conventions. It shapes the agent's approach but never overrides the safety principles or the permission engine.- Benchmark harness β
larb bench <suite>reports resolution rate and cost-per-task (the Β§14 metrics). - Interactive session β bare
larbopens a persistent conversation: follow-ups carry history, Esc interrupts a run (resumable, steerable), typing mid-run queues the next task, and/plandrafts a reviewed plan before executing. - CLI, TUI, and a headless editor bridge (
larb bridge, a stdio JSON protocol) β true incremental streaming, diff review, approval prompts, live cost meter.
packages/
governors/ trust Β· permission Β· cost Β· audit
providers/ model provider abstraction + adapters
sandbox/ scoped command execution
context/ repo map + markdown memory
core/ orchestrator loop + capability tools
cli/ CLI + TUI
Requires Node.js β₯ 20 and pnpm.
pnpm install
pnpm build # bundles the CLI to packages/cli/dist/index.js
export ANTHROPIC_API_KEY=sk-ant-... # or any provider key (see Providers)
# Run from the repo:
pnpm larb ask "What does the orchestrator loop do?" # read-only
pnpm larb run "Add a --version flag to the CLI" # prompts for trust + each write/exec
pnpm larb audit # audit log + costpnpm build
npm link ./packages/cli # puts `larb` on your PATH (resolves deps from this checkout)
larb versionA published npm i -g @larb/cli is planned but not yet available β use npm link
for now.
See config.example.toml for configuration.
Larb is model-agnostic. No provider is privileged in the codebase β each one is a row in a preset table that bundles its wire transport, base URL, key env var, default models, and pricing. You own your keys (read from the environment, never from repo config) and your routing.
Select a provider with kind in ~/.larb/config.toml and export its API key:
[provider]
kind = "deepseek" # see the table below
# apiKeyEnv = "..." # override the key env var (optional)
# baseURL = "..." # point at any compatible endpoint (trusted config only)
[provider.models] # optional β omit to use the preset's defaults
orchestrator = "deepseek-chat" # strong model: plans & orchestrates
worker = "deepseek-chat" # cheap/fast model: subtasks & compactionexport DEEPSEEK_API_KEY=...kind |
Provider | API key env var | Transport |
|---|---|---|---|
anthropic |
Anthropic Claude | ANTHROPIC_API_KEY |
Anthropic Messages |
openai |
OpenAI GPT | OPENAI_API_KEY |
OpenAI Chat |
ollama |
Local (Ollama) | β (no key, no spend) | Ollama |
deepseek |
DeepSeek | DEEPSEEK_API_KEY |
OpenAI-compatible |
gemini |
Google Gemini | GEMINI_API_KEY |
OpenAI-compatible |
groq |
Groq | GROQ_API_KEY |
OpenAI-compatible |
mistral |
Mistral | MISTRAL_API_KEY |
OpenAI-compatible |
xai |
xAI Grok | XAI_API_KEY |
OpenAI-compatible |
openrouter |
OpenRouter | OPENROUTER_API_KEY |
OpenAI-compatible |
together |
Together AI | TOGETHER_API_KEY |
OpenAI-compatible |
perplexity |
Perplexity | PERPLEXITY_API_KEY |
OpenAI-compatible |
Most providers expose an OpenAI-compatible Chat Completions API, so they share a
single audited adapter β adding a new one is a new table row, not new code. For
any endpoint not listed, set kind = "openai" (or "anthropic") and override
baseURL + apiKeyEnv.
Routing. The strong orchestrator model plans and drives the loop; a cheap
worker model handles delegated subtasks and context compaction, so long runs
stay inexpensive. Both are per-provider and overridable.
List providers and check which keys are set from the CLI:
larb providers # table of all providers + whether each key is set
larb providers deepseek # base URL, default models, and config snippetThis is an alpha. Be aware of these before relying on it:
- Sandbox isolation depends on a container runtime. Real isolation (the
"safe to open an untrusted repo" property) needs docker or podman
installed. Without one, Larb falls back to a reduced-isolation host
subprocess β cwd-scoped with host secrets stripped, but the host filesystem
and network are reachable. The active level is printed at the start of every
run; verify the container path with
docs/verify-container.md. - Container egress allow-listing filters proxy-respecting clients (package managers, curl, fetch); airtight raw-socket blocking awaits a microVM backend.
- SWE-bench ships as a loader + grading primitives; full graded runs need the
dataset repos and per-repo test commands (
docs/swebench.md). - The editor bridge (
larb bridge) has a VS Code extension MVP ineditors/vscode(build from source; not yet on the marketplace). - Windows is CI-verified (typecheck, tests, build, and binary smoke run on the ubuntu/macos/windows matrix) but hasn't had the same live-model run verification as macOS/Linux yet.
- MCP support is stdio-only for now (the common case); SSE/HTTP transports and MCP resources/prompts are planned behind the same client interface. MCP servers are global-config-only β a repo can never define one.
larb benchruns tasks autonomously with auto-approval β point it at a disposable checkout only.- Prompt caching and streaming are covered by tests against mocked transports; confirm savings/behavior against your live provider.
Apache-2.0. Contributions require signing the CLA. The "Larb" name
and logo are trademarks and are not granted by the code license β see
TRADEMARK.md. See SECURITY.md for
coordinated disclosure, CHANGELOG.md for release notes, and
threat-model.md for the attack classes Larb designs out.
