Skip to content
the-open-enginePublic

About

Enforced QA for Claude Code, Codex and Copilot. Draw the graph: code review, your own reviewers and test agents, any model per node. Slop goes back to repair, and the agent that wrote the code can't skip it.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2.0k stars

Watchers

3 watching

Forks

Latest commit

 

History

669 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

zeroshot: the agent that wrote the code should not be the one that says it works. Independent review and repair.

 

Join the zeroshot community on Discord Website · zeroshot.sh The Open Engine · theopenengine.com X · @OpenEngineHQ LinkedIn

Release npm Build Opcore Coverage Docs License: MIT

 

Starred by engineers at Google, Meta, Shopify, Uber, Booking.com, GitHub, Cloudflare, Atlassian and JetBrains

zeroshot

The agent that writes the code should not be the one that decides it works.

We built zeroshot because we were tired of being gaslit by agents telling us broken code was ready.

zeroshot runs a QA process your coding agent can't skip. You define it as a graph: which agents review, what must pass, where a rejection loops back. In the built-in software-change workflow, one agent implements, independent agents review, and failures go back to a bounded repair loop. Git delivery waits for both reviewers to accept. zeroshot doesn't replace Claude Code, Codex or Copilot. It runs one of them as the worker and as the reviewers.

Use the built-in graph or draw your own: code review, your own reviewers and test agents, any model per node. Save the setup as a profile for the next task.

Skills and review graphs

A skill is advice: the agent still decides whether to follow its reviewer's feedback and when to stop. A zeroshot graph defines the required reviews and repair loop before execution; the engine follows it mechanically until the reviews accept or the retry limit is reached.

Add custom agent nodes: a security reviewer, a bugfinder that writes tests, or an agent that checks the app in a browser. Choose each node's model and instructions, which nodes run in parallel, and where rejected work goes for repair.

The product site, zeroshot.sh, has the FAQ.

▶ Video: a single agent vs. zeroshot (scripted)

Questions, ideas, or a run worth showing? Join the zeroshot community on Discord.

Get started

Prefer Cloud? Cloud setup →

Run locally

npm install -g @the-open-engine-company/zeroshot

The installer requires Node.js 18 or newer and installs a verified native binary for Linux x64/arm64, macOS x64/arm64, or Windows x64. It also installs one zeroshot skill for Codex, GitHub Copilot, and Claude Code at user scope.

For local execution, install and sign in to Codex, Claude Code, or GitHub Copilot. Local runs can reuse the harness's existing login, including subscription-backed sessions. See the installation guide for harness prerequisites.

Run your first task

This example uses Codex and keeps delivery local. The worker edits the current Git worktree, so start in a clean worktree intended for the task.

Create input.json, replacing the example task with a change appropriate for your repository:

{
  "task": "Add JSON output to the status command and cover it with focused tests."
}

Create runtime.json. Replace YOUR_MODEL_ID with a model identifier supported by your installed Codex CLI:

{
  "harness": "codex",
  "provider": "openai",
  "model": "YOUR_MODEL_ID",
  "effort": "high"
}

Start the run. Add --validate-only to check the graph, runtime configuration, and input first:

zeroshot run \
  --title "Add JSON status output" \
  --template software-change \
  --input input.json \
  --uniform-runtime-config runtime.json

Open another terminal to inspect the run:

zeroshot ui

Visit http://127.0.0.1:4173/ui/ and open Runs. Stopping the UI server leaves active runs running. The CLI also provides zeroshot list, zeroshot status RUN_ID, and zeroshot logs RUN_ID.

For other harnesses and providers, see Runtimes and connections.

Run in the cloud

Connect GitHub and bring your own model API keys. Cloud setup guide →

What the built-in workflow does

  1. A worker implements the task.
  2. Acceptance and code reviewers check the result independently, in parallel.
  3. Rejected work goes to one repair worker with both reviews' feedback, then through both reviews again.
  4. With delivery enabled, accepted work proceeds through the configured Git and CI steps.
  5. Delivery conflicts return through repair and review.

The graph defines the sequence, parallel steps, retry paths, and exit conditions before execution starts. Runs are bounded, and events are recorded in a durable SQLite ledger. A passing run means its configured checks accepted the work; coverage depends on the requirements, reviewers, tests, and environment you provide.

Animated zeroshot workflow: implementation, parallel acceptance and code review, repair, and optional Git delivery
Both reviews run again after a repair. Git delivery is optional.
Static diagram

Auto-research

The built-in auto-research template runs ten iterations by default; set a positive options.iterations value in the run input to choose another count. Three scouts propose directions, then a planner chooses an experiment and either the current candidate or a restorable archived candidate as its starting point. Independent judges check the evidence and method, while the progress judge decides whether the result should replace the current candidate under the task charter. A valid result can go into the archive without replacing it. An auditor checks the ledger, archive, and workspace before the next iteration.

Inspect the graph with zeroshot template show auto-research.

Bring your own graph topology

Choose each agent's model and instructions, which steps run in parallel, and when to retry. Example topology, adding stages after the review loop:

Implementation
      ↓
Code review + acceptance review
      ↓
Bugfinder: adversarial tests
      ↓
E2E tests
      ↓
Delivery

The software-change review loop, auto-research graph, and delivery are built in. You define other stages, their tools, services and credentials, and where failures go for repair.

Inspect the built-in graph as a starting point:

zeroshot template show software-change

See the graph contract for custom graph authoring and Prepare a runtime environment for dependency and service setup.

Save and reuse a profile

Each step can use its own model. Pass --runtime-config with one binding per node in place of --uniform-runtime-config:

{
  "harness": "claude",
  "provider": "anthropic",
  "size": "medium",
  "nodes": {
    "worker": { "kind": "agent", "model": "WORKER_MODEL_ID", "effort": "high" },
    "acceptance": { "kind": "agent", "model": "REVIEW_MODEL_ID", "effort": "high" },
    "code": { "kind": "agent", "model": "REVIEW_MODEL_ID", "effort": "max" },
    "review_repair": { "kind": "agent", "model": "WORKER_MODEL_ID", "effort": "high" }
  }
}

Runs with delivery enabled also need a delivery_repair binding. Run zeroshot template show software-change to list the node names. See the RuntimePlan reference for every field.

A profile stores a graph and its runtime settings. Use Profiles in the browser UI to edit and save your configuration. Keep different profiles for different kinds of work, or reuse one for the next task.

Once you've saved a local profile named my-profile, run another task with it:

zeroshot run \
  --title "My next task" \
  --profile local:my-profile \
  --input input.json

The CLI and UI share the same local profiles. See UI setup.

Choose delivery and execution

Keep the first result local, then enable the delivery mode that fits your process:

Delivery Behavior
No delivery flag Keep the work local
--push Push the managed branch
--pr Prepare a mergeable pull request without merging
--ship Proceed through PR, CI, and merge

PR and ship runs address visible pull request review feedback by default; pass --no-pr-feedback to ignore it.

Graphs can execute locally or on a self-hosted target. Each environment needs its own tools and authentication setup.

Self-hosted: run the Docker target

Keep execution and durable state on infrastructure you control. The target image includes the native engine plus pinned Codex, Claude, and GitHub Copilot harness CLIs.

docker run --detach --restart unless-stopped --name zeroshot-target \
  -p 127.0.0.1:8080:8080 \
  -v zeroshot-data:/var/lib/zeroshot \
  ghcr.io/the-open-engine/zeroshot-target:latest

zeroshot target add local --url http://127.0.0.1:8080 --direct

The target also serves its profile editor and run viewer at http://127.0.0.1:8080/ui/.

See the target image guide for persistent storage, network isolation, builds, and HTTPS.

A failed run with a recoverable workspace can be restarted or resumed.

Research graph

zeroshot also includes an auto-research graph for ten bounded iterations of experiments with independent review of evidence, method, and progress. It keeps adopted, record-only, and aborted results under .zeroshot/research; add --push when each finalized iteration should be published.

zeroshot template list
zeroshot template show auto-research

See Execution for the research workflow's decision and evidence rules.

Community

  • Discord: ask questions, share runs and graphs, and talk to the team.
  • GitHub Issues: reproducible bugs and feature requests.

Reference

Also check out: Opcore.

Development

npm ci
npm run check
cargo test --workspace # Unix; Windows: powershell -NoProfile -File scripts/test-windows.ps1

Node.js builds the static UI and supports repository tooling and npm delivery; Rust serves the UI. See UI development, CONTRIBUTING.md, PUBLISHING.md, and SECURITY.md.

License

MIT. See LICENSE.

About

Enforced QA for Claude Code, Codex and Copilot. Draw the graph: code review, your own reviewers and test agents, any model per node. Slop goes back to repair, and the agent that wrote the code can't skip it.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2.0k stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages