A fast, single-binary agentic CLI for coding and automation, powered by open-weight models through OpenRouter. Written in Go.
Its defining feature is a cost-ladder router: per task class, open-agent tries the cheapest adequate model first and escalates to a pricier one only when the learned pass-rate proves the cheap one inadequate — so routine work runs on free/cheap models and only genuinely hard tasks reach expensive ones. It learns which model is "minimally sufficient" for each kind of task from an execution-grounded verification gate, and shows you the whole thing in a built-in dashboard.
- One-shot coding/chat:
open-agent code "add a --json flag and update the tests". - Autonomous orchestration —
open-agent do "<goal>"decomposes a goal into a dependency-aware task DAG, runs the tasks across parallel multi-model workers, and gates each code task through an execution-first verifier (its acceptance command must exit 0; failures feed back a bounded Reflexion retry).--plan-modelpins a cheap orchestrator while the ladder still picks the best worker per task. - Hardened sandbox —
open-agent sandbox run "<task>"gives the agent root in a throwaway, host-isolated container for real automation (provisioning, package installs, running services, autonomous bugfixes on cloned repos). Named multi-environment sessions, selectable egress (--net full|none|https). - Observability —
open-agent dashboardserves a local web UI over~/.open-agent: runs, costs, event traces, per-model stats, prompt-cache hit rates, and the live router ladder.
# Homebrew (macOS arm64 / Linux amd64) — recommended
brew install imhassla/tap/open-agent # updated on each tagged release# or: GitHub Release binaries (darwin_arm64, linux_amd64)
ARCH=darwin_arm64 # or: linux_amd64
gh release download -R imhassla/open-agent -p "*${ARCH}*" -D /tmp/oa
tar -xzf /tmp/oa/open-agent_*_${ARCH}.tar.gz -C /tmp/oa
sudo install -m 755 /tmp/oa/open-agent_*/open-agent /usr/local/bin/open-agent
# or build from source
make install # builds + installs (single binary)
# or: go install github.com/imhassla/open-agent@latest
make install-treesitter # optional: richer multi-language repo_map (CGo tree-sitter)Releases are opt-in: put [release] (patch), [release:minor], or [release:major] in a commit message on main and CI tags + builds + publishes it and auto-bumps the Homebrew tap. A manual git tag vX.Y.Z && git push also works: git tag v0.1.0 && git push origin v0.1.0.
open-agent needs an OpenRouter API key (get one). Set it once, machine-wide:
mkdir -p ~/.config/open-agent
echo 'OPENROUTER_KEY=sk-or-...' > ~/.config/open-agent/.env
chmod 600 ~/.config/open-agent/.env
open-agent ask "reply with exactly: OK" # verifyResolution order (first match wins): the OPENROUTER_KEY env var → a .env in the
working directory → ~/.config/open-agent/.env. For a single shell session you can just
export OPENROUTER_KEY=sk-or-... instead. See .env.example.
An interactive session is saved per-project to .open-agent/session.json in the current
directory. Resume it (across restarts, updates, or machines — copy the dir) with:
open-agent --continue # or: open-agent chat --continue / -cGlobal learning (memory, router ratings, do run traces) lives in ~/.open-agent/.
An interactive session snapshots the working tree before every turn into a shadow git
repo under ~/.open-agent/checkpoints/ — it never touches your real git history, and it
captures untracked files and shell side effects. Undo a turn's changes:
/rewind # list checkpoints
/rewind 3 # restore files AND conversation to before turn 3
/rewind 3 code # restore only the files (keep the conversation)
/rewind 3 chat # restore only the conversation (keep the files)
So you can "try something risky, then rewind if it fails." Disabled automatically inside a home/system dir or when a nested git repo is present.
open-agent # interactive session (intent auto-detected)
open-agent code "…" # one-shot coding task in the cwd
open-agent ask "…" # one-shot chat, no tools
open-agent do "…" # plan a goal into a DAG and run it
open-agent do --plan-model minimax/minimax-m2.1 "…" # cheap orchestrator, ladder workers
open-agent models # list families and their per-role models
open-agent dashboard # local observability UI (default :8787)
open-agent sandbox run "…" # run a task as root in a hardened containerFlags: -m/--model <slug> pin a model · -f/--family <name> pick a family (cold-start
prior) · --plan-model <slug> pin the orchestrator only · --json one machine-readable
result envelope on stdout · --max-cost/--max-tokens/--deadline budget ceilings ·
--no-route disable the dynamic router · --sandbox run bash/verification in Docker.
Unknown flags are a hard error (never sent to the model as prose).
Best-of-N for code — open-agent code --candidates 3 "…" runs N (2–4) candidate
workers in parallel, each in an isolated throwaway checkout of HEAD and each pinned
to a different model family (default rotation qwen,glm,minimax; override with
--families a,b). Each candidate tree is verified (go build ./... && go test ./...
when it has a go.mod; otherwise the worker's own ok verdict is trusted), then the
single best diff is git apply'd onto your real tree — winner = verified success with
the fewest changed lines, ties broken by cost. Requires a git-clean tree (the only
resulting dirt is the winning diff, reviewable with git diff). --max-cost is split
evenly across the candidates, so total spend stays under the same cap. If every
candidate fails, nothing is applied and the exit code is 1. --sandbox forwards to
the candidates — each subprocess mounts its own isolated tree in Docker, so
containerized bash composes cleanly with best-of-N isolation.
For scripted/agent callers (e.g. a supervising LLM delegating subtasks), see
AGENTS.md — the machine contract (--json envelope, cost caps, tier
policy, sandbox recipes).
Guardrails (on by default): mutating file tools refuse to write outside the working
directory (absolute paths out of tree, ../ escapes, symlink targets), and bash rejects
a tight list of catastrophic command shapes (recursive rm of / or ~, force-push —
including --force-with-lease: an unattended worker rewrites no history, fork bomb,
mkfs/dd to raw devices, curl-pipe-to-shell). Scope honesty: the bash layer filters
known-catastrophic shapes only — an arbitrary bash redirect can still write outside the
tree; the hardened Docker sandbox is the boundary for genuinely untrusted work. A denial
is a normal tool error the model can read and adapt to. OPEN_AGENT_NO_GUARDRAILS=1
disables both layers for legitimate out-of-tree work.
Guardrails are user-extendable: put deny rules in ~/.config/open-agent/guardrails
(global) and/or ./.open-agent/guardrails (project-local), one per line —
name: regex denies matching bash commands (RE2), write name: glob denies writes to
matching paths even inside the tree (filepath.Match globs, matched against the
relative path AND its basename, so write no-env: *.env covers any depth; matching is
case-folded). Write rules bind the file tools only — a bash redirect is out of their
reach, so pair a write rule with a bash regex rule when that matters. Best-of-N applies
the winning diff through the parent's rules, so project-local write rules hold there too.
User rules add to the built-ins and can never remove them; a malformed line or
invalid regex is warned to stderr and skipped, never fatal.
Related switches: OPEN_AGENT_ADAPTIVE=0 turns off the per-model lab-recommended
prompting/sampling layer; OPEN_AGENT_NO_REASONING_REPLAY=1 turns off MiniMax
interleaved-thinking replay.
Every role's candidate set is all families' models for that role, ordered by observed
per-task cost (list price for cold rungs). For each
(role, task-class) bucket the router:
- tries the cheapest unproven rung first, stops at the cheapest proven-reliable rung (EWMA pass-rate ≥ 0.5), skips proven-unreliable ones, and re-probes a long-benched cheap rung occasionally so a bad day isn't a permanent ban;
- learns from an execution-grounded gate (a code task's acceptance command must exit
0), from one-shot outcomes, and from whole
do-run success; - escalates mid-run: a failed verification is recorded before the retry re-picks, and a budget-pressure downshift trims unaffordable rungs when the run is low on budget.
Watch it live in open-agent dashboard → Router ladder.
| Command | Purpose |
|---|---|
code |
autonomous coding agent (cwd) |
ask |
chat that can web-search when a question needs current facts |
research |
read-only web research (grounded, cited search) |
do |
plan → parallel multi-model DAG → verify |
improve |
review → fix → verify cycle (fixes uncommitted) |
schedule |
recurring jobs: add/list/remove + run daemon |
sandbox |
run as root in a hardened box |
dashboard |
local observability web UI |
models / runs / replay / bench |
introspection & self-eval |
Model families: kimi, glm, google, grok, deepseek, qwen, minimax, mistral
(the family is the router's cold-start prior; the ladder learns from there). :free
catalog models are excluded from the ladder (field-tested too weak for real work) but
can still be pinned explicitly with -m <slug>:free.
bash · read_file · write_file · edit_file · go_replace_func (AST-based
whole-function replacement — no exact-text matching) · todo_write (maintain an
in-task step plan for long tasks) · glob · grep · repo_map ·
web_search (grounded + cited, via OpenRouter Sonar) · web_fetch · memory_store · memory_retrieve ·
final_answer. Code workers also get spawn_subagent (one-level delegation),
read_artifact, and code_consensus (best-of-N with a cross-family judge).
By default bash runs on the host with your privileges. Only run open-agent on
inputs/goals you trust, or isolate it:
--sandboxrunsbashand verification in a network-isolated, resource-capped Docker container.open-agent sandbox …runs the whole agent as root inside a hardened, throwaway container that is isolated from the host: no--privileged, no host bind mounts, no Docker socket, no host namespaces, default seccomp/AppArmor kept, escape-prone capabilities dropped, pids/memory/cpu capped, ephemeral by default. Egress is selectable (--net full|none|https). The OpenRouter key is injected at run time, never baked into the image.
Do not point the agent at untrusted repositories or prompts without the sandbox.
Not zero-dependency: the default build uses go-git (Apache-2.0), yaegi (Apache-2.0),
and go-diff (MIT); the optional treesitter build adds go-tree-sitter (MIT). All
permissive and compatible with this project's MIT license (see LICENSE).
main.go CLI: arg parsing, subcommand dispatch, system prompts
internal/config OpenRouter key resolution (env / .env)
internal/llm OpenRouter client (Chat + ChatStream), model slugs, pricing
internal/agent ReAct loop + tool registry; parallel dispatch
internal/tools bash, file ops, web, repo_map, git baseline/verify
internal/rating persistent per-(role,model) pass-rate/cost store (the ladder's memory)
internal/orchestrator role→model cost-ladder router; planner→DAG; verifier; scheduler
internal/budget atomic shared run budget (steps/tokens/cost/wall)
internal/sandbox hardened Docker sandbox (posture, multi-env, egress modes)
internal/dash local observability web server + embedded UI
internal/telemetry JSONL run log + failure-pattern hints
The agent loop is a ReAct cycle: call the model with the tool schemas, dispatch returned
tool calls concurrently, append results, and repeat until a direct answer or final_answer.
agent.Doer is an interface, so the whole loop runs offline against a scripted fake in
tests.
make test # go test ./...
make race # go test -race ./...The planner-prompt quality has a live A/B eval harness gated behind an env var:
OPEN_AGENT_LIVE_EVAL=1 go test ./internal/orchestrator/ -run TestPlanPromptABLive.