LLM-powered local code review for git diffs. Run a "second pair of eyes" review on your staged changes, a specific file, a single commit, or a PR-style three-dot range — without leaving the terminal.
Pick a provider once, then:
commitbrief # review your staged changes
commitbrief diff HEAD # review working tree vs HEAD
commitbrief diff main...feature/x # review a PR
commitbrief --unstaged --dir app/Models # narrow any scope to a directoryOutput is rendered as colored markdown in the terminal, plain markdown to a file, or strict JSON for tooling — your choice.
A real reviewer is the gold standard, but they aren't always available the moment you stage a change. CommitBrief gives you a quick, structured read on your diff before another human (or your future self) sees it.
- Local-first. Diffs and review output stay on your machine. The only network egress is to the provider you chose.
- Provider-agnostic. Anthropic, OpenAI, Gemini, or Ollama as
API-backed providers;
claude-cli,gemini-cli, andcodex-clireuse your local Claude Code / Gemini / Codex CLI subscription (no extra API key). - Cache aware. Re-running on an unchanged diff is essentially free —
one disk read, no token spend.
--verboseshows what you saved. - Custom review rules. A repo's
COMMITBRIEF.mdis sent as the system prompt; per-userOUTPUT.mdcontrols how findings are formatted.
CommitBrief ships an eval harness (make eval) that scores real review
output against a 23-fixture known-answer corpus — 23 planted defects
across security, correctness, concurrency, resource-leak, error-handling
and performance categories, plus 3 clean controls a good review must stay
silent on. About a quarter of the corpus is a held-out slice that
prompt and corpus tuning never inspect, so each cell below reports
dev · held — the tunable slice and the held-out generalization slice
separately (ADR-0018). Numbers are from make eval-live, captured
2026-05-29 (mean of Runs live runs each):
| Model | Recall (dev · held) | FP-rate (dev · held) | Precision (dev · held) | Runs |
|---|---|---|---|---|
| Claude Haiku 4.5 | 1.00 · 1.00 | 0.00 · 0.00 | 0.70 · 0.62 | 5 |
| Claude Sonnet 4.6 | 1.00 · 1.00 | 0.00 · 0.50 | 0.68 · 0.48 | 3 |
| Claude Opus 4.8 | 0.94 · 1.00 | 0.00 · 0.00 | 0.61 · 0.53 | 3 |
| Gemini 2.5 Flash | 0.96 · 1.00 | 0.44 · 0.00 | 0.84 · 0.56 | 3 |
| OpenAI GPT-4o | 0.85 · 1.00 | 0.44 · 0.33 | 0.79 · 0.75 | 3 |
- Recall — share of planted defects caught. Every model recalls the full held-out slice; the dev dips (Opus, GPT-4o) come from the harder multi-finding dev fixtures, not from missing whole defects.
- FP-rate — findings landing on a clean-control line (flagging a benign change). Note where the noise lives: Sonnet trips the held-out clean control; Gemini and GPT-4o trip the dev ones.
- Precision — a conservative floor: any finding outside the answer key counts as a false positive, but on these small diffs many "extra" findings are legitimate secondary observations (a second panic, an ignored error) rather than noise. The terser models (GPT-4o, Gemini) score higher precisely because they say less — at the cost of recall. Read recall + FP-rate as the cleaner signals; precision is sensitive to how exhaustively the corpus is annotated.
The two slices are not difficulty-matched — the split exists to catch
overfitting in future tuning (a dev gain that doesn't carry to held-out),
not for a direct dev-vs-held comparison today. Reproduce any row with
COMMITBRIEF_EVAL_PROVIDER=<name> make eval-live, which prints FULL / DEV /
HELD-OUT scorecards (using the key already in ~/.commitbrief/config.yml).
brew install CommitBrief/tap/commitbriefscoop bucket add commitbrief https://github.com/CommitBrief/scoop-bucket
scoop install commitbriefgo install github.com/CommitBrief/commitbrief/cmd/commitbrief@latestPre-built binaries for Linux, macOS, and Windows on amd64 and arm64 are attached to each tagged release at github.com/CommitBrief/commitbrief/releases.
commitbrief upgrade # check, confirm, install
commitbrief upgrade --check # report only; install nothingupgrade detects how the binary was installed and does the right thing for
it: Homebrew, Scoop and go install are handed to their own package
manager, while a manually installed binary is downloaded from GitHub
Releases, SHA-256 verified against the release checksums, and swapped in
place. If the binary's directory is not writable, nothing is downloaded and
the exact command you need is printed — CommitBrief never runs sudo
itself. Only the binary is replaced; bundled man pages are not installed.
This is the only network request CommitBrief makes on its own behalf, and only when you run this command. There is no automatic update check.
The v1.0.0 line is an API freeze. CLI flag surface, the JSON
schema v1 ({schema, content, findings, summary, meta} — emitted by
--json), COMMITBRIEF.md and OUTPUT.md formats, and the public
config keys all follow strict semver from v1.0.0 onwards — breaking
changes ship in v2.x. The current line is v1.0.0-rc.1, the freeze
checkpoint; anything locked in here is the long-term contract.
Upgrading from v0.x? See the migration guide in
CHANGELOG.md — the scope
flags (--commit / --branch / --pull-request) and --yes
semantics changed during the v0.9.x line.
# 1. One-time setup: pick a provider, paste your API key, run a ping.
commitbrief setup
# 2. (optional) Write a project-specific review rules file.
commitbrief init
# 3. Stage some changes and review them.
git add path/to/changed.go
commitbrief --stagedcommitbrief list prints the full command reference; commitbrief dry-run --staged walks the pipeline without spending tokens.
Default TTY output is a framed view: a header, a status line, one
bordered panel per finding (colored by severity), and a one-line
summary footer. Findings are ordered critical → info.
commitbrief v0.6.0 · provider: anthropic/claude-sonnet-4-6 · cache: miss
analyzing 3 files · 42 added · 11 removed · COMMITBRIEF.md loaded
┌─ CRITICAL ─ internal/auth/session.go:142 ──────────────────────────────────┐
│ SQL fragment built from request input │
│ │
│ String concatenation feeds db.Query() directly, bypassing the prepared │
│ statement path used elsewhere in this package. │
│ │
│ - q := "SELECT * FROM sessions WHERE token = '" + tok + "'" │
│ + q := "SELECT * FROM sessions WHERE token = $1" │
│ rows, err := db.Query(ctx, q, tok) │
└─────────────────────────────────────────────────────────────────────────────┘
┌─ HIGH ─ internal/db/migrate.go:73 ─────────────────────────────────────────┐
│ NOT NULL column added without a default │
│ │
│ The new column has no DEFAULT, so the migration will fail on any table │
│ that already has rows. Either backfill in a prior migration or add a │
│ DEFAULT before the constraint. │
└─────────────────────────────────────────────────────────────────────────────┘
┌─ LOW ─ internal/api/handler.go:201 ────────────────────────────────────────┐
│ Wrapped error duplicated in message │
│ │
│ The format string already contains "%w"; the prefix repeats the wrapped │
│ error verbatim, producing "auth failed: auth failed: …" in logs. │
└─────────────────────────────────────────────────────────────────────────────┘
✓ Done in 4.2s · 3 findings · 8421 tokens · Cost: $0.0319
Five severity levels — critical, high, medium, low, info —
colored red/orange/yellow/blue/grey. info items are always shown;
suppress them with a user-side OUTPUT.md template (see Configuration).
Re-run the same command on the same diff and the footer switches to
Saved: $0.0319 — a local cache hit, no provider round-trip.
--json emits the raw findings (documented schema), --markdown runs
your OUTPUT.md template against the findings and writes plain text
suitable for >> review.md.
# Review scopes
commitbrief # = --staged (default scope)
commitbrief --unstaged # working-tree changes
commitbrief diff HEAD # working tree vs HEAD (git-diff passthrough)
commitbrief diff HEAD~3 HEAD # the last three commits
commitbrief diff main feature # one branch vs another
commitbrief diff main...feature # three-dot PR-style range
# Narrow any scope with repeatable path filters (exact paths or globs)
commitbrief --unstaged --file app/Http/Controllers/API.php --file routes/web.php
commitbrief --unstaged --dir database/seeder --dir app/Models
commitbrief diff HEAD~3 HEAD --dir docs
commitbrief --staged --file '*.go' # gitignore-style glob (any depth)
commitbrief --staged --file 'internal/**/*.ts' # anchored recursive glob
commitbrief --staged --exclude-file '*_test.go' # denylist; wins over the includes
commitbrief --staged --dir internal --exclude-dir internal/cli
# Select the commits themselves — author, date window, message or branch name.
# Any of these walks history instead of reading the index, so they replace
# --staged/--unstaged rather than combining with them.
commitbrief --author alice --author bob # either person's commits
commitbrief --author alice@example.com # name or email, case-insensitive
commitbrief --start-date 2026-01-01 # on or after (inclusive)
commitbrief --end-date 2026-03-31 # on or before (inclusive)
commitbrief --text payment # commit message OR branch name
commitbrief --committer carol --merges # committer identity; keep merges
commitbrief --author alice --start-date 2026-06-01 --dir internal # all combinable
commitbrief diff main..develop --author alice # bound the walk to a range
# Plain-language change digest (read-only; no findings)
commitbrief summary # what's staged, grouped by area
commitbrief summary main...develop # a range; uses the commit messages in it
commitbrief summary HEAD~3 HEAD -o NOTES.md # write the digest to a file
# Commit message (writes to git, with confirmation)
commitbrief commit # suggest a message for the staged diff, then commit
commitbrief commit --type conventional # pick a format (-t); see "commitbrief commit" below
commitbrief commit --generate 3 # offer 3 alternatives to choose from (-g)
commitbrief commit --yes # commit the first suggestion non-interactively
# Setup and rules
commitbrief setup [--local] # provider + API key wizard
commitbrief setup --alias[=cbr] # install a shell alias for commitbrief (bash/zsh/fish/PowerShell/cmd)
commitbrief providers list|use|test # switch active provider without re-running setup
commitbrief config show|get|set # inspect / tweak the merged YAML config
commitbrief init [--force] # write COMMITBRIEF.md + OUTPUT.md template
commitbrief compress [--level=balanced] [--dry-run] # shrink COMMITBRIEF.md (preview first if you want)
commitbrief doctor # health-check the pipeline
commitbrief install-hook [--hook=...] # install a git hook that runs commitbrief
commitbrief upgrade [--check] # check GitHub Releases and install a newer CommitBrief
commitbrief dry-run # pipeline preview; no API call
commitbrief list # command reference
commitbrief mcp # run an MCP server over stdio (agent review gate; see "MCP server")
commitbrief guard # gate a review against .commitbrief/policy.yml (see "Policy gate")
# Cache maintenance
commitbrief cache clear # wipe every cached LLM response for this repo
commitbrief cache prune [flags] # bounded cleanup; defaults --keep-last 500 --older-than 7d
commitbrief cache stats # entry count, size, age range, per-provider breakdown
commitbrief cache inspect <key> # one entry's metadata (add --show-content for the body)Global flags: --json, --markdown, --output <file>, --copy,
--suggest-commit (after the review, suggest a Conventional Commit
message for the staged diff; prints to stdout, requires --staged, not
with --json/--markdown/--output), --compact, --no-cache,
--fail-on=<sev>, --min-severity=<sev>
(hide findings below this severity in the rendered output; --json and
--fail-on still see the full set), -f/--file (repeatable; exact
path or gitignore-style glob), -d/--dir (repeatable; exact prefix or
glob), --yes, --verbose, --quiet, --lang,
--provider, --model, --cli <claude|gemini|codex> (shorthand for the
CLI-tool-backed providers; mutually exclusive with --json /
--markdown), --with-context (CLI providers only — let the host CLI
read project files beyond the diff to ground the review; see below),
--allow-secrets (acknowledge a flagged credential in
the diff), --no-cost-check (skip cost preflight),
--show-prompt (print the exact system + user prompt that would be sent,
then exit — no provider call, no cost; honours --output), --no-flaky
(skip the flaky-test detector below), --sandbox-rerun[=N] (opt-in
sandbox-rerun confirmation of flagged flaky tests; see below),
--no-architecture (skip
architecture-aware review; see below), --update-baseline /
--no-baseline (signal-control baseline; see below), --color. See
commitbrief --help.
Before the model is called, a static pre-pass scans the added lines of any changed test files for high-precision flakiness anti-patterns:
- hard-coded sleeps / fixed waits (
medium):time.Sleep,Thread.sleep,Task.Delay,asyncio.sleep,*.waitForTimeout, numericcy.wait,usleep,sleep(<n>); - unseeded randomness (
low):Math.random, Pythonrandom.*, Gomath/rand; - brittle selectors (
low, js/ts): absolute/positional XPath, CSS:nth-child/:nth-of-type, Cypress.eq(<n>), Playwright.nth(<n>)— stabledata-testid/ role / attribute selectors are never flagged; - over-mocking (
low): a single test function that sets up more than five mocks/stubs (jest.mock/spyOn,when(…).thenReturn,sinon.stub,patch(…),gomock/.EXPECT(), Mockery) — pinned to implementation, not behaviour; - time-dependent assertions (
low): a wall-clock read (time.Now(),Date.now(),new Date(),datetime.now,System.currentTimeMillis()) used directly in an assertion instead of an injected clock.
Matches merge into the normal findings, so they render, count toward
--fail-on, and --copy like any other finding — but they are deterministic
and reproducible: no model call, no JSON-schema change. On by default for the
API/mock providers; turn it off per-run with --no-flaky or persistently with
review.flaky: false. CLI-tool-backed plain-text providers are unaffected for
now.
Sandbox-rerun confirmation (opt-in). The rules above infer flakiness from anti-patterns; sandbox-rerun confirms it by actually re-running a flagged test in isolation N times and classifying it by the observed pass/fail mix:
- mixed pass + fail → confirmed flaky (the finding is kept and its suggestion notes the empirical confirmation);
- all fail → a real failure, not flakiness — the test is genuinely red, so the note says so plainly (don't quarantine it as a flake);
- all pass → transient / resolved — the flake did not reproduce, so the
finding is demoted to
infoand won't trip a commit-stage--fail-on.
Confirmation requires a double opt-in: a positive --sandbox-rerun[=N] /
review.sandbox_rerun (bare flag uses N=5) and a non-empty
review.sandbox_command — either alone stays a no-op. sandbox_command is a
list of argv elements (never a shell string), each rendered as a Go
text/template over {{.File}} (repo-relative), {{.Line}}, and {{.Test}}
(the enclosing test function name), then executed directly with
exec.CommandContext — no shell is invoked:
review:
sandbox_rerun: 5
sandbox_command: ["go", "test", "-count=1", "-run", "^{{.Test}}$", "./..."]config set review.sandbox_command is rejected — hand-edit the config file
directly (config get prints it read-only). Each rerun attempt is bounded by
a 2-minute timeout (a hung test costs one attempt, not the whole review), and
a stderr notice names the configured command template — the un-rendered
argv, e.g. -run ^{{.Test}}$ — once per review, before any per-test rendering
happens; this is the review path's first code-execution stage, so it is never
silent. The command runs against the working tree, not the staged
snapshot a review may be scoped to. commitbrief mcp and commitbrief guard
never run it, unconditionally, regardless of flag or config — an agent host
must not execute repository code unattended (ADR-0033 §6); the static
findings still return, unconfirmed.
Sandbox-rerun is not cached. The flaky pre-pass (and any sandbox-rerun
confirmation inside it) runs before the cache lookup, so its findings can
merge into both a cache hit and a fresh call. That means a cache hit does
not skip the rerun: re-running the same command against the same diff
still executes the configured command again, even though the review body
itself is served from cache and the footer reports Saved. Total cost scales
with the number of flagged findings — worst case findings × N × 2 minutes,
with no cap and no progress output while it runs.
Test-name resolution is Go-only. {{.Test}} resolves via go/parser
against *_test.go files; every other language resolves to no name, so that
finding skips the rerun (with a stderr warning) and keeps its bare static
signal. Python/JS/PHP/Java tests keep full static flaky detection — they just
never get sandbox-rerun confirmation. See ADR-0033 for the full execution-boundary
rationale.
Resolution also skips any file the Go toolchain would not compile on this
machine: a //go:build constraint that is not satisfied, a _windows_test.go
on Linux, anything under testdata/. That includes build tags your
sandbox_command itself supplies — a -tags=integration in your command does
not make an //go:build integration file resolvable, so those tests keep their
static finding and skip confirmation. The alternative would be naming a test
go test cannot select, which exits 0 and would report a flake as "did not
reproduce" without ever running it.
Three layers keep the noise down. --min-severity (above) is display-only.
The other two are true removals — a removed finding no longer counts toward
--fail-on, no longer appears in --json findings[], and is hidden from the
rendered output:
-
Baseline (per-developer, gitignored). On a brownfield repo, run
commitbrief --update-baselineonce to accept the current findings — it writes their fingerprints to.commitbrief/baseline.jsonand does not filter that run. Every later run then surfaces only new findings; the known ones are removed. The fingerprint is resilient to line drift (a finding that moves up or down the file stays baselined). The file is never committed — it lives under the already-gitignored.commitbrief/, so CI and a reviewer's gate apply no baseline and see everything (it can't be used to hide a real bug). Re-accept any time with--update-baseline; ignore the baseline for one run with--no-baseline; disable persistently withreview.baseline: false. -
Inline suppression (in committed source). Silence one finding with a visible reason by adding a comment on the offending line — or the line directly above it:
result := db.Query(userInput) // commitbrief-ignore[high]: input is parameterized below
Use
commitbrief-ignore: <reason>to silence any finding on the line, orcommitbrief-ignore[<severity>]: <reason>to silence only that severity. The comment prefix doesn't matter (//,#,--,/* */all work). Because the marker is in committed source, a reviewer sees it in the diff.
Neither filter is silent: optional meta.baselined / meta.suppressed counts
appear in --json (the schema stays v1) and a one-line N baselined · M suppressed footer prints to stderr.
If your repo ships an architecture.json — the config of the sibling tool
archlint, which deterministically
lints import-boundary rules — CommitBrief reads it and feeds a compact summary of
your declared layers and their allowed / forbidden import edges into the
review prompt. The reviewer can then flag a diff that crosses an architectural
boundary, e.g. an import that adds a domain → db dependency the architecture
forbids:
{
"layers": { "domain": ["internal/domain"], "db": ["internal/db"] },
"rules": { "domain": [], "db": ["domain"] }
}(rules lists, per layer, the layers it is allowed to import; [] means it
may import no other layer.) This is a strict one-way read of archlint's public
config — CommitBrief never lints the import graph or enforces anything itself
(run archlint check in CI for the deterministic gate); it only grounds the LLM
so it can reason about the change. A missing or malformed file is a transparent
no-op that never breaks a review.
On by default; turn it off per-run with --no-architecture or persistently with
review.architecture: false. Point it at a non-default location with
review.architecture_file. Because the architecture summary folds into the
prompt, editing architecture.json automatically invalidates stale cached
reviews, while a repo without the file keeps a byte-identical cache key.
Already using pre-commit? Add CommitBrief to your
.pre-commit-config.yaml:
repos:
- repo: https://github.com/CommitBrief/commitbrief
rev: v1.8.0
hooks:
- id: commitbrief # language: golang — pre-commit builds the binary
# - id: commitbrief-system # or: use an already-installed binary on PATHBoth ids run the review gate on the staged diff (--staged --fail-on=high,
override args: to change the gate). This is distinct from commitbrief install-hook (below), which writes a native git hook script and needs no
framework — pick whichever fits your setup. The hook runs non-interactively, so a
flagged secret aborts the commit (it never auto-confirms).
Generate a commit message from the staged diff and, after you confirm,
run git commit. This is the only command that writes to git — every
review path is read-only.
commitbrief commit # suggest one message, confirm (default Yes), commit
commitbrief commit -t conventional+body # conventional subject + a generated body
commitbrief commit -g 4 # pick from 4 alternatives
commitbrief commit --provider openai --model gpt-5.4-mini
commitbrief commit --yes # CI/non-interactive: commit the first suggestion--type/-t— message format:plain(default),conventional,conventional+body,gitmoji,subject+body.--generate/-g <N>— produce N alternatives (1–10) and choose one in an arrow-key selector. A single provider call generates all N.--provider/--model/--cli— select the backend, same as a review. Messages are always written in English regardless of--lang.- Defaults come from the
commit.typeandcommit.generateconfig keys when the flags are omitted (precedence: flag > config > built-in). - The pre-send
.commitbrief/**guard, secret scan, and cost preflight run on the staged diff before the call; the suggestion is cached. - With nothing staged it errors (stage with
git addfirst). On a non-TTY without--yesit errors, because it cannot show the confirm or selector.--yescommits the first suggestion (it does not bypass the secret scan or cost preflight).
The tool never auto-stages and never edits files — it only runs
git commiton changes you already staged, and only after you say Yes.
Explain a set of changes in plain language — a short, human-readable digest of what changed (and, when the commit messages make it clear, why), grouped by logical area rather than file by file. Read-only; it produces no findings.
commitbrief summary # digest the staged diff
commitbrief summary --unstaged # digest unstaged working-tree changes
commitbrief summary main...develop # digest a range (git-diff passthrough, like `diff`)
commitbrief summary HEAD~3 HEAD # digest the last three commits
commitbrief summary main...develop -o RELEASE.md # write the digest to a file
commitbrief summary main...develop --lang tr # Turkish digest
commitbrief summary --cli claude # use a host CLI tool as the backend
commitbrief summary --cli claude --with-context # let the CLI agent read beyond the diffExample output:
Invoice Service: Rounding bug in fee calculation fixed. (a1b2c3d)
Auth: Token refresh flow added. (d4e5f6a)
DB: Index added to the invoices table. (f7a8b9c)
- Scope mirrors the review surface: no args ⇒ staged (default),
--unstagedfor the working tree, or positionalgit diffarguments for an arbitrary range — exactly like thediffsubcommand.--file/--dirnarrow it further. - For a range, the commit messages in that range are taken into account and each line is attributed to the short commit hash(es) responsible. Staged / unstaged changes have no commits, so their lines carry no attribution.
- Output is plain text;
-o/--outputwrites it to a file.--langis honoured (e.g.--lang tr). - Provider selection is identical to a review:
--provider/--model, or--cli claude|gemini|codexto use a host CLI tool. With a CLI provider,--with-contextlets the agent read files beyond the diff to ground the digest (it errors on an API provider). - Reuses the pre-send
.commitbrief/**guard, secret scan, and cost preflight, and is cached. Emits no findings, so--json,--markdown,--suggest-commit,--fail-on, and--min-severityare rejected. Never writes to git.
By default a review sees only the diff. With --with-context, a
CLI-backed provider (--cli claude|gemini|codex) is allowed to read
other files in the repo — callers of the changed code, type definitions,
sibling modules, project conventions — to ground its review in the wider
codebase. The diff stays the subject of the review; the rest is context.
The host CLI runs read-only (it never modifies your tree) in the
repository root. API providers can't read files, so the flag errors for
them.
⚠ Security: with
--with-contextthe agent decides which files to read, so file contents beyond the diff — including untracked secrets (.env, key files) — can reach the provider's backend. The pre-send secret scan covers the diff only, not files the agent reads on its own. CommitBrief prints this caution on every--with-contextrun. Use it on repositories you trust.
Four API providers + two CLI-tool-backed providers ship in the box:
| Provider | Models | Notes |
|---|---|---|
| Anthropic | Claude Opus 4.8 (default), Sonnet 4.6, Haiku 4.5 | Ephemeral prompt caching (5 m TTL) cuts repeated input cost ~10×. Opus 4.8 advertises a 1 M-token context. |
| OpenAI | GPT-5.4-mini (default), GPT-5.5, GPT-5.5-pro, GPT-4o, GPT-4o-mini | Automatic prompt caching at ≥1024-token prefixes. gpt-5.5-pro runs via the Responses API (not Chat Completions) and can take several minutes per review. |
| Google Gemini | Gemini 3.5 Flash (default), 3.1 Pro, 3.1 Flash-Lite | ~1 M-token context windows. gemini-3.1-pro-preview is a preview model. |
| DeepSeek | deepseek-chat, deepseek-reasoner | OpenAI-compatible API (DEEPSEEK_API_KEY); JSON is prompt-driven (degrades gracefully). |
| Mistral | Mistral Large / Small, Codestral | OpenAI-compatible API (MISTRAL_API_KEY). |
| Cohere | Command R+ / R, Command A | Cohere's OpenAI-compatibility endpoint (COHERE_API_KEY). |
| Ollama | Whatever you've ollama pull'd |
Local-only, no API key, no per-token cost. |
claude-cli |
Whatever your local Claude Code uses | Subprocess of claude -p - — no API key on our side; reuses your Claude Code subscription. commitbrief --cli claude --staged. |
gemini-cli |
Whatever your local Gemini CLI uses | Subprocess of gemini -p — no API key on our side; reuses your Gemini CLI auth. commitbrief --cli gemini --staged. |
codex-cli |
Whatever your local Codex CLI uses | Subprocess of codex exec --sandbox read-only --skip-git-repo-check — no API key on our side; reuses your Codex CLI (ChatGPT) auth. commitbrief --cli codex --staged. |
CLI-backed providers emit pre-formatted plain text — they bypass the
structured-findings JSON path, the per-finding cards renderer, and the
--fail-on severity gate (the host CLI's response shape isn't our
contract to enforce). The review block is bracketed top and bottom
with a -------------------- rule (the same separator used between
findings) and written to stdout, so
commitbrief --cli claude --output review.md writes the file just
like the API providers do; --json / --markdown are rejected
upfront. Useful when you've already paid for a Claude or Gemini CLI
subscription and don't want to manage a second API key.
Adding a provider is one new package under internal/provider/<name>/.
The
remote prsubcommand (below) requires an API provider when it posts to GitHub —claude-cli/gemini-cli/codex-clidon't produce structured findings to anchor comments. (In--no-postmode it only prints locally, so CLI providers work there.)
commitbrief remote pr <ID> reviews a GitHub pull request and writes the
result back to GitHub: each finding becomes an inline review comment and
the review is submitted with a verdict (approve / comment /
request-changes). It drives your local gh CLI — no hosted bot, no extra
auth.
commitbrief remote pr 42 # PR #42 in the current repo
commitbrief remote pr CommitBrief/web#10 # cross-repo (owner/repo#N)
commitbrief remote pr 42 --request-changes-on=high
commitbrief remote pr 42 --no-post # review locally, write nothing to GitHub
commitbrief remote pr 42 --no-post --output review.md # …or --json / --cli gemini, etc.--request-changes-on=<critical|high|medium|low> opts in to a
request-changes verdict at or above that severity. Without the flag the
verdict is never request-changes — a clean or info-only PR is approved,
anything else is left as a plain comment. Inline comments are posted
either way. --repo owner/repo overrides git-context repo discovery.
Requires an API provider. --fail-on is ignored here — the GitHub verdict
replaces the exit-code gate.
--no-post turns remote pr into a read-only review: it fetches the
PR diff via gh and renders the result to your terminal exactly like a
local review, writing nothing to GitHub (no comments, no verdict).
Because the output is local, the flags posting mode rejects all apply —
--json, --markdown, --output, --copy, --compact, --cli, and
--fail-on — and there's no self-PR restriction (you can review your own
PR). Results are cached like any local review. Handy for triaging a PR,
piping findings into another tool, or reviewing with a CLI provider.
Each comment is anchored to the diff side its line lives on — RIGHT
(new file) for added/context lines, LEFT (old file) for removed ones.
A finding whose line falls outside the diff (or whose POST is rejected)
is not dropped: it is appended to the review summary so nothing is lost.
Run CommitBrief on pull requests with the CommitBrief Review GitHub Action:
# .github/workflows/commitbrief.yml
on: pull_request
permissions:
contents: read
pull-requests: write # comment mode posts the review
jobs:
review:
runs-on: ubuntu-latest
steps:
- uses: CommitBrief/commitbrief-action@v1
with:
provider: anthropic
api-key: ${{ secrets.ANTHROPIC_API_KEY }}It posts each finding as an inline review comment plus a verdict
(comment mode, via remote pr), or runs an exit-code gate
(mode: gate, via diff --fail-on). You can also drive the binary
directly in any workflow: commitbrief diff <base>...<head> --fail-on=high.
commitbrief mcp runs a Model Context Protocol server over stdio so an
AI agent or host (Claude Desktop, an agent runtime, an MCP-aware IDE) can call
CommitBrief as a tool — typically a self-review gate the agent runs before it
submits code. It speaks JSON-RPC 2.0 over the MCP stdio transport, is
stdlib-only (no MCP SDK, no new dependency), and is fully opt-in: it changes
nothing about the existing commands.
The server exposes one tool, review, which runs the exact same review
pipeline as commitbrief --json (diff acquisition, filtering, the pre-send
guard + secret scan, cost preflight, cache, the flaky-test pre-pass, and signal
control) and returns the structured findings (JSON schema v1) plus a short
text summary. It does not re-implement the review — it reuses runReview.
Tool: review — all arguments optional:
| Argument | Type | Meaning |
|---|---|---|
staged |
bool | review the staged diff (default) |
unstaged |
bool | review the working tree (⊥ staged) |
diff |
string[] | git diff range args, e.g. ["HEAD~3","HEAD"] or ["main...feature"] |
provider |
string | override the configured provider |
model |
string | override the configured model |
fail_on |
string | critical|high|medium|low|info|any|none — reported as a gate failure in the summary; findings are still returned |
min_severity |
string | hide findings below this severity in the returned set |
no_flaky |
bool | skip the deterministic flaky-test detector |
The result carries two content blocks: a one-line summary (finding counts,
provider, and a GATE FAILED note when --fail-on trips) and the schema-v1
JSON document. A genuine failure (no repo/changes, provider error, or an aborted
secret-scan guard) comes back as an MCP tool error.
Wiring it into a host. Register commitbrief mcp as a stdio MCP server. For
a Claude Desktop-style host config:
{
"mcpServers": {
"commitbrief": {
"command": "commitbrief",
"args": ["mcp"]
}
}
}The host launches the process, performs the initialize handshake, discovers the
review tool via tools/list, and calls it via tools/call. The server reads
requests on stdin and writes responses on stdout until the host closes the stream;
diagnostics go to stderr. See the MCP server wiki page
for the full handshake and a worked example.
commitbrief guard is a declarative merge gate: it evaluates a review's
actionable findings against a .commitbrief/policy.yml and exits non-zero when
the policy is breached. It is richer than the single --fail-on=<severity>
threshold — a per-severity budget — and is aimed at gating high-volume (often
AI-authored) pull requests.
Create the policy (the gate is opt-in — no file means no gate):
# .commitbrief/policy.yml
version: 1
thresholds: # max findings allowed per severity (omit or ~ = unlimited)
critical: 0
high: 0
medium: 5
low: ~
total: 20 # optional overall capSharing the policy with your team. CommitBrief adds
.commitbrief/to your.gitignorethe first time it writes a cache entry, which also ignores the policy file. To version-control the gate, add a negation below that line:.commitbrief/ !.commitbrief/policy.ymlThe baseline (
.commitbrief/baseline.json) is deliberately not shared — it is per-developer, so it can never hide a finding from CI or a reviewer.
Then gate a change — two modes:
# run-mode: review the diff, then evaluate (reuses the full pipeline)
commitbrief guard # staged diff
commitbrief guard --unstaged
commitbrief guard --diff main...HEAD
# consume-mode: evaluate a review you already produced — no provider call
commitbrief --json --staged > review.json
commitbrief guard --from-json review.jsonIt evaluates the findings that survive baseline + suppression (signal
control) — exactly what --json shows. The exit code is 0 (pass) or
non-zero (blocked); a load/parse failure also blocks (a merge gate must not
pass when it cannot prove the change is within policy). --json emits a machine
verdict ({passed, counts, total, violations}). guard complements --fail-on
(the simple one-threshold gate) — use either or both. Rule-id-scoped allow/deny
lists are not yet supported (findings carry no stable rule id). See
the guard wiki page.
Two-tier YAML config with field-level merge:
- User:
~/.commitbrief/config.yml— defaults that apply everywhere - Repo:
./.commitbrief/config.yml— overrides for this repo (gitignored by default; runcommitbrief setup --localto write it)
Plus environment variables for credentials and runtime tweaks, and
CLI flags for one-off overrides (--provider gemini --model gemini-3.5-flash).
| Variable | Effect |
|---|---|
ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY |
Provider credentials. Overrides the matching providers.<name>.api_key in config. |
OLLAMA_HOST |
Sets providers.ollama.base_url when not set in config. |
COMMITBRIEF_PROVIDER |
Selects the active provider (same as --provider / config.provider). |
COMMITBRIEF_MODEL |
Overrides the active provider's model. |
COMMITBRIEF_CONFIG |
Absolute path to the user-level config file; replaces the default ~/.commitbrief/config.yml lookup. Useful for tests and ephemeral CI environments. |
COMMITBRIEF_NO_COLOR, NO_COLOR |
Force ANSI color off (overrides --color always). |
LANG |
No longer drives language (ADR-0021): language is config-driven (--lang → repo output.lang → user output.lang → English). |
# ~/.commitbrief/config.yml
version: 1
provider: anthropic # default provider
providers:
anthropic:
model: claude-opus-4-8
pricing: # optional: override built-in $/1M rates
claude-opus-4-8: # (cost preflight / verbose footer / cache)
input_per_1m: 5.0
output_per_1m: 25.0 # omitted fields keep the built-in value
openai:
model: gpt-5.4-mini
ollama:
model: qwen2.5-coder:14b
base_url: http://localhost:11434
output:
lang: en # AI output language (any recognized lang, e.g. fr); UI localizes for en/tr only
stream: true
color: auto # auto | always | never
cache:
enabled: true
ttl_days: 7
max_size_mb: 0 # 0 = unlimited; >0 evicts oldest entries past the cap
guard:
secret_scan: true # scan diff + rules for credential patterns before sending
token_preflight: false # opt-in: confirm/abort when the prompt overflows the model's context window
injection_scan: true # warn (never abort) if a non-default COMMITBRIEF.md/OUTPUT.md has prompt-injection phrasing
secret_patterns: # additive user credential regexes; built-ins always run (ADR-0024)
- name: "Internal Service Token"
regex: 'INT-[0-9]{10}'
command:
default: "" # args applied to a bare `commitbrief`; empty = `--staged`
commit:
type: plain # default --type for `commitbrief commit` (plain|conventional|conventional+body|gitmoji|subject+body)
generate: 1 # default --generate (number of message alternatives)
review:
flaky: true # deterministic flaky-test detector pre-pass (ADR-0022); --no-flaky overrides per-run
sandbox_rerun: 0 # opt-in sandbox-rerun confirmation (ADR-0022): re-run a flagged test N times in isolation; 0 = off; --sandbox-rerun[=N] overrides per-run
sandbox_command: [] # the rerun executor (ADR-0033): argv list, e.g. ["go", "test", "-run", "^{{.Test}}$", "./..."]; requires sandbox_rerun > 0 too (double opt-in); hand-edit only, config set rejects it; Go-only test-name resolution
baseline: true # apply the user-private signal-control baseline (ADR-0027); --no-baseline overrides per-run, --update-baseline rewrites it
architecture: true # architecture-aware review (ADR-0030): read architecture.json into the prompt; --no-architecture overrides per-run
architecture_file: "" # override the architecture.json discovery path (relative to repo root, or absolute); empty = auto-discoverA bare commitbrief reviews staged changes (commitbrief --staged). To
change that default, set command.default to the argument string you'd
otherwise type:
command:
default: --unstaged --cli gemini # now `commitbrief` == `commitbrief --unstaged --cli gemini`It applies only to the truly bare invocation. The moment you pass any
flag or subcommand — commitbrief --staged, commitbrief --json,
commitbrief dry-run — the default is bypassed and you get exactly what
you typed. Empty/unset keeps the built-in --staged. Tokens are
whitespace-split (no shell quoting).
Review content lives in two files:
COMMITBRIEF.mdat the repo root — team-shared review rules, perspectives, project context. Sent to the LLM as the system prompt. Committed to git..commitbrief/OUTPUT.md(or~/.commitbrief/OUTPUT.md) — per-user Gotext/templateapplied locally to the findings for--markdownand--output <file>.md. Never sent to the LLM. The template has access to.Findings(typed[]Finding{Severity, File, Line, Title, Description, Language, Snippet}) plus helpers likegroupBySeverity,upper,countFiles. Gitignored.
commitbrief init writes both templates from the embedded defaults.
Two independent axes: which commits are reviewed, and which files within them.
--author, --committer, --start-date, --end-date and --text select a
set of commits. Setting any of them switches the scope from "the index" to a
history walk, so they cannot be combined with --staged / --unstaged — those
have no commits yet. commitbrief commit and commitbrief remote pr reject
them outright (the first describes the index; the second reads its diff from
gh, not local git).
| Flag | Matches |
|---|---|
--author |
author name or email, case-insensitive substring; repeatable, OR'd |
--committer |
committer name or email; repeatable, OR'd |
--start-date YYYY-MM-DD |
author date on or after this day (inclusive) |
--end-date YYYY-MM-DD |
author date on or before this day (inclusive — unlike git's bare --until) |
--text |
the commit message, plus commits unique to a branch whose name contains the text |
--max-commits N |
cap the selection (default 200); truncation is always reported |
--merges |
keep merge commits, which are excluded by default |
Different kinds are AND'd, multiple values of one kind are OR'd:
--author alice --author bob --start-date 2026-06-01 means "(Alice or Bob)
and since June".
The revision range walked is HEAD by default, or the range you give a
subcommand: commitbrief diff main..develop --author alice. The resulting
diff is the concatenation of the matching commits' patches, not a
cumulative range diff — so a file changed in three of them appears three
times, and no unmatched commit's work leaks in.
Branch-name matching is best-effort by nature: a squash- or rebase-merged branch no longer owns its commits, so nothing will be found for it.
Three ignore layers, applied in order. Later layers win, so a !pattern in
.commitbriefignore can revert a built-in exclusion:
- Built-in defaults — binaries, lock files,
vendor/**,node_modules/**, generated code, build artifacts, IDE/OS noise. .commitbriefignoreat the repo root — gitignore syntax, team-shared.COMMITBRIEF.mdsemantic filter — natural-language rules the LLM applies to whatever survives the first two layers.
On top of those, --file / --dir narrow to a path allowlist and
--exclude-file / --exclude-dir remove from it. All four share the same
matching rules (exact path or gitignore-style glob), and exclusion is applied
last, so it always wins.
commitbrief dry-run reports how many commits matched and how many files each
layer removed.
Requires Go 1.25+.
git clone https://github.com/CommitBrief/commitbrief
cd commitbrief
make build # → ./commitbrief (ldflags inject version/commit/date)
make test # unit + integration tests (live providers skipped)
make lint # golangci-lint v2
make smoke # end-to-end pipeline check; no API key needed
make bench # diff pipeline + cache hit benchmarks
make manpage # regenerate man/*.1 from cobra
make test-live # provider tests against real APIs (keys required)
make license-check # GPL-3.0 compatibility auditmake help lists everything.
Does CommitBrief replace human review? No. It's a first pass — a fast sanity check before a teammate (or your future self) looks. The default rules deliberately target low-risk, high-signal stuff: obvious bugs, missing nil checks, accidental secrets. Treat output as a checklist to skim, not a verdict.
Where does my code go?
Diffs leave your machine only when sent to the provider you picked.
Anthropic, OpenAI, and Gemini get the diff + your COMMITBRIEF.md over
HTTPS to their official endpoints. Ollama is local-only; nothing leaves
the host. Review output is rendered locally and cached under
./.commitbrief/cache/ — never uploaded.
Will it break my workflow if the LLM provider is down?
The CLI fails loudly and exits non-zero. There's no degraded mode that
silently skips review. Use commitbrief dry-run to test the pipeline
end-to-end without an API call.
How do I exclude generated code or vendored files?
Built-in defaults already skip vendor/**, node_modules/**, lock
files, binaries, and most generated artifacts. Drop a
.commitbriefignore at the repo root for project-specific rules
(gitignore syntax, supports !negation to revert a built-in).
commitbrief dry-run --staged reports how many files each layer
removed.
When does the cache invalidate? The cache key is a SHA-256 of `diff + system prompt + provider + model
- lang + schema version
. Change any of those and you get a fresh review. Default TTL is 7 days; configurable viacache.ttl_days. Setcache.max_size_mb(>0) to bound the on-disk cache: writes that push it past the limit evict the oldest entries first (the entry just written is never evicted). Inspect it withcache stats/cache inspect `.
Can I run it in CI?
The primary target is the developer's terminal, but the CI-friendly
pieces are in place: --fail-on=<severity> (or --fail-on=any)
returns a non-zero exit code when a finding meets or exceeds the
threshold, --json emits the structured-findings document machine-
readably, and commitbrief install-hook scaffolds a pre-commit /
commit-msg / pre-push hook locally. For pull-request CI there's the
CommitBrief Review GitHub Action
(see "Continuous integration" above), or you can drive the binary
directly from any workflow.
Why GPL-3.0?
The CLI is end-user software, and copyleft keeps forks and
redistributions open. New dependencies must stay GPL-compatible
(MIT/Apache-2.0/BSD/ISC/MPL-2.0/LGPL-3.0+ are fine);
make license-check enforces it.
GPL-3.0-or-later. Provider SDK dependencies are Apache-2.0 or
MIT; the full audit is make license-check.
See CONTRIBUTING.md for the project-specific build
and test flow, and the
org-wide CONTRIBUTING guide
for inbound-equals-outbound licensing and PR conventions.
Bug reports and questions are welcome in Issues and Discussions.