Skip to content

Latest commit

 

History

History
2608 lines (2251 loc) · 231 KB

File metadata and controls

2608 lines (2251 loc) · 231 KB

CLI reference

The full set of o2b verbs. After o2b install-cli they are on PATH; the local checkout can also be used without installing the symlinks by running commands through scripts/o2b and scripts/vault-log.

Most verbs accept --vault <path> and --config <path>; the values default to the active profile from o2b init. JSON output is available via --json on every read verb. Commands that do not own a semantic JSON contract return a redacted fallback envelope when --json is passed.

Core

o2b status                    Show config / vault status
o2b init                      Bootstrap the vault profile (idempotent)
o2b init --interactive        Guided first-time setup wizard
o2b install                   Runtime installer: bare form detects, --target <name> plans, --target <name> --apply writes, --check verifies (see "Exit codes" below); --friction reports what each host costs instead of acting on one
o2b install-cli               Symlink o2b and vault-log into ~/.local/bin
o2b-mcp                       Console-script alias for `o2b mcp`, forwards all flags
o2b doctor                    Run vault + adapter checks
o2b index                     Rebuild the Markdown page index
o2b export-config             Write a redacted config snapshot
o2b secrets list|status       Inspect $secret:NAME references without printing values; `--vault <path>` joins that vault's custody store (names and availability only, never a decrypt; without the flag both commands are byte-identical to the store-less report)
o2b mcp                       Run the MCP tool server (stdio by default; --transport http binds loopback, and a credential - a key, --api-key or, kept out of the process list, OPEN_SECOND_BRAIN_MCP_API_KEY, or a non-empty per-agent token map from `o2b mcp token mint` - is required only for a non-loopback --host); a token-matched request authenticates as that token's agent for the request, the shared key keeps the process identity, and `mcp_tokens_required` / OPEN_SECOND_BRAIN_MCP_TOKENS_REQUIRED (default false) refuses credential-less requests while a non-empty token map exists; --scope full|writer|catalog, --tool-profile full|writer|catalog|recall|minimal (an unknown profile exits 2 rather than serving the full surface), --host-target <runtime>, --harness <id>, --probe, --allow-tool, --disable-tool, --max-tools
o2b state status|migrate|rollback
                              Inventory the state this vault holds, move it to another directory, or put it back (see "State surfaces" below)
o2b tool-call                 Invoke an MCP tool handler from the CLI
o2b version                   Print the installed Open Second Brain version; --version is a synonym
o2b help --json               Print the command/flag manifest as JSON
o2b completions --shell zsh   Print completions for bash|zsh|fish|elvish|nushell|powershell
o2b uninstall                 Print uninstall plan; --apply-local cleans config; --remove-cli removes symlinks
o2b update                    Update Open Second Brain across all detected runtimes; --target <name> / --dry-run / --force / --json
o2b bootstrap                 One-command harness provisioning for --target <name> (codex, grok, opencode run the adapter's idempotent apply; generic prints the payload plus manual steps; claude-code and zcode are plugin verify-only). --token mints mcp_token_<target> and prints the material exactly once - never on argv, never in a harness config; a second identical run is a byte-identical no-op; --rotate re-mints under the same name (effective on the next request, no restart); --check verifies drift; receipt at <vault>/.open-second-brain/bootstrap.lock.json
o2b mcp token                 Per-agent MCP token management over the vault's hash-at-rest store: mint (derives mcp_token_<agent> unless --name), rotate, revoke, list; mint and rotate print the material exactly once with a shown-once notice, list shows metadata only

o2b version (since v1.56.0)

o2b version prints the installed version and exits 0:

$ o2b version
1.56.0
$ o2b version --json
{"version":"1.56.0"}

o2b --version is a synonym and renders the same line, --json included - the flag is routed to the verb rather than printing a second line that could drift from it. Both are declared in the command manifest, so shell completions and o2b help --json offer them.

The verb takes no positional argument. o2b version latest is a usage error rather than a print of the local version: nothing in this build learns what any other version is, and answering a question about the latest release with the installed one would be the wrong answer wearing the right shape. The MCP handshake has always carried the same fact as serverInfo.version, and both now read one constant.

o2b install exit codes

Every code the verb can return, in one table. A CI script that gates on o2b install --check should branch on these rather than on the printed table:

Code Meaning
0 Success, or --check found no drift. A target the operator never installed reports not-installed and still exits 0
1 I/O or runtime error during --apply
2 Usage error: unknown --target, a bad --format value, or a vault that is not configured
3 --check found drift in at least one target
4 --apply hit a user-modified managed block; re-run with --force to overwrite it
5 --check found a runtime it proved unreachable (since v1.46.0)

Code 5 is a behaviour change, and a breaking one for any script that treated a zero exit as "everything is fine". Before v1.46.0 a mcp-unreachable verdict shared exit 0 with not-installed, so o2b install --check reported success for a runtime whose own adapter had just failed to reach it - the one verify state backed by a live query rather than a file read. The two answers no longer share a code.

Drift outranks unreachable when both are present, and the order is the operator's repair order rather than a severity ranking: a runtime can be unreachable BECAUSE its config drifted, so --apply is the first move and a liveness verdict taken over a wrong config means nothing.

mcp-unreachable is reachable only from an adapter that asks a runtime CLI whether the servers are registered. Two targets can produce it: copilot-cli (copilot mcp list) and, since v1.50.0, codex (codex mcp list) - see install/copilot-cli.md and install/codex.md. Those are exactly the two rows of src/core/runtime/host-facts.ts that declare a hostProbe; every other adapter verifies off disk, so it can report ok, drift, or not-installed but never mcp-unreachable. What a refuting probe MEANS differs between the two install modes - see "Host probes and what a verify proves" below.

o2b doctor exit codes

Code Meaning
0 Every check passed, and with --readiness every probe answered
1 At least one check FAILED, or with --readiness at least one probe proved its surface broken
6 With --readiness: no check failed, and at least one probe could not find out

Code 6 is a behaviour change for --readiness runs and is deliberately the number o2b search check already spends on a probe that did not complete; a test asserts the two cannot drift. A probe that exceeded its per-check budget used to be counted as a failure and exit 1, so the same healthy machine exited 0 when idle and 1 when loaded - the verdict was a property of the load average rather than of the installation. Such a probe now reports unknown with the elapsed budget as its reason, and an unmeasured surface is not folded into the 0 that would claim it was checked either. A proved failure outranks an unmeasured probe, so 1 never hides behind 6.

The --json shape gains readiness_summary (probes, failed, unknown) beside the existing readiness array whenever --readiness is passed, so a caller reads the same three-way answer the exit code carries; ok is read off that exit code, and being two-valued it means only "not established as healthy" when false.

o2b doctor selftest (since v1.64.0)

Plain o2b doctor proves configuration, and two of its checks write probe files, but none opens the search store - so a machine can pass every check and still be unable to init, migrate, index, search or survive concurrent writes. The selftest verb drives a real throwaway store end to end on temporary storage only: open and migrate, an index pass, a document upsert, keyword and trigram search, a concurrent-write roundtrip through the real writer lock, delete, close - stopping at the first failure, each stage reported as a doctor-check-shaped entry that names the fault and the fix. The harness never opens the configured vault (its store is pinned inside the temp tree it removes on every exit path) and never reaches a provider (the semantic lane is forced off; o2b search check answers that question). --json carries no timestamps, paths or durations, so two healthy runs on one machine stringify byte-identically. Exit codes reuse the doctor table above: 0 every stage passed, 1 a stage failed, 6 the harness itself could not run (temp storage unavailable), because it then established nothing about the store.

The automatic Brain upgrade check

o2b doctor reports self_heal_upgrade. After a plugin update the full-scope o2b mcp server and the SessionStart hook bring a stale _brain.yaml / _BRAIN.md current in a detached worker (o2b brain upgrade --self-heal, not an operator flag). A failed attempt is recorded per device in .open-second-brain/self-heal-upgrade.json, and the check then fails with the error, the time of the last attempt, the number of consecutive failures and the time from which the next automatic attempt is due; its fix is o2b brain upgrade --dry-run. The same record appears in o2b brain status as the self-heal-upgrade-failed problem and at the end of o2b brain upgrade --dry-run / --check output (JSON: self_heal_failure).

The next automatic attempt waits a cooldown of one hour, doubling with each consecutive failure up to 24 hours, instead of repeating the same failing upgrade on every start. o2b brain upgrade --apply --yes ignores the cooldown and clears the record once nothing is pending.

Since v1.73.0 an upgrade applies exactly the plan it showed. o2b brain upgrade --apply and the automatic worker plan once and apply that plan; each planned row records the bytes it was read with, or that the file was absent. A managed file that changed on disk after the plan was read is refused before the snapshot is taken: nothing is written, and the message names the files and advises o2b brain upgrade --dry-run to review a plan that reads the edit. Each write checks the file again right before it replaces it, so an edit that lands mid-apply is left as it is; when earlier files were already rewritten the message names them and advises the same dry-run re-plan; a rollback is deliberately not offered, because the pre-apply snapshot predates the rewritten files, so its own drift guard would refuse it, and forcing it would destroy the edit. With --apply --json a refusal prints { "ok": false, "error", "run_id", "drifted" }, where drifted lists the refused paths. A file that holds bytes which are not valid UTF-8 and was not edited still upgrades.

The codegraph partner check

o2b doctor consults the optional codegraph partner when its CLI is on PATH and the working directory is a code project, by spawning codegraph status -j <repo> once per discovered project. That costs about 0.7 s against a warm HOME and about 5 s against a cold one, because the partner caches under HOME.

partner_codegraph_disabled: "true" in the config file, or OPEN_SECOND_BRAIN_PARTNER_CODEGRAPH_DISABLED=true in the environment, turns that consultation off. Default off, so the doctor behaves exactly as before unless the switch is set. It is a config key rather than a flag because the fact it records ("the partner is installed on this machine and asking it costs seconds") is a property of a machine, and because the three callers of the doctor include the MCP vault_health tool and the OpenClaw extension, which take no flags. o2b partner codegraph report is unaffected: an answer the operator asked for directly is never suppressed by a switch about background cost.

A doctor run with the switch on still prints a code_graph line saying which switch silenced it, so a check that did not run is distinguishable from one that ran and passed. A machine with no codegraph CLI, and a directory that is not a code project, still both print nothing at all - those two remain indistinguishable from each other in the doctor's output.

o2b install --friction (since v1.50.0)

o2b install --friction                       Every registered runtime, one row per capability dimension
o2b install --friction --target <t>          One runtime
o2b install --friction --base <a> --compare <b>
                                             Only the dimensions the two answer differently

A REPORT, and it exits 0 whatever it finds: the matrix has no notion of a bad answer, only of an honest one. The only non-zero exits are usage refusals - an unknown runtime name (the message lists the registered ones), --target combined with the diff pair, or a diff missing half of its pair. --json emits the same values under a friction key, and the diff under friction_diff.

Six dimensions, computed from src/core/runtime/host-facts.ts and from the live adapter rather than from a table maintained beside this page:

Dimension What the cell answers
install-mechanism the step kinds the adapter's own plan uses - json-merge, managed-block, subprocess, file-copy, symlink, print - with the artifact paths it writes on the citation line
tool-ceiling the host's published per-workspace tool limit: declared: N tools with its source, unbounded with its source, or unknown with the reason nobody has one
tool-profile the profile the generated registration will actually carry, and which tier of the ladder produced it; not carried for a target that writes no MCP command line at all
verify-evidence what a --check on this target is evidence OF: host probe: <command> for a runtime that can be asked, configuration comparison only for one that cannot
session-transcripts the transcript roots this adapter DECLARES, resolved against the machine, with each root's glob and on-disk format. Declared, not measured - whether the directory exists here is a discovery question
session-parser which adapter in this build reads those roots, or none ships naming the format nothing parses

A diff prints only the differing dimensions and COUNTS the rest, because a diff that reprints everything is the table the reader already asked to skip; the count is what stops the silence reading as "there was nothing else to compare".

Three dimensions were considered and deliberately left out rather than guessed at, and src/core/install/friction.ts records why: whether the host takes MCP at all (no structural fact distinguishes a managed block that registers servers from one that does not), whether lifecycle hooks are wired (one adapter knows, privately), and whether runtime identity is stamped (applied at apply time, and no read-only adapter surface carries it). Each becomes derivable the day the adapter contract states it.

A cell whose input cannot be resolved reads unresolved and carries the obstacle. That is how a Brain/_brain.yaml this run cannot parse arrives: the ladder below refuses rather than defaulting, and a report that turned that refusal into a stack trace would stop being total.

Host probes and what a verify proves (since v1.50.0)

An adapter whose runtime publishes a way to ask it now asks. The two declared probes are codex mcp list and copilot mcp list; both are keyless, start no model turn, and write nothing. A clean --check on those targets prints what the host said - beside what the config file declares where an artifact backs the registration, and instead of it where the host's own registry is the only record.

The blanket note configuration comparison; no MCP handshake attempted survives only where no probe is declared. Where one IS declared and could not run, the check says which obstacle it hit - so a skip is never rendered as a handshake:

host probe skipped: `codex` is not on PATH
host probe skipped: `codex mcp list` exited 1: <the command's first line of stderr>

The wait is capped at 10 seconds. A host CLI that blocks on a network call or an interactive authentication prompt would otherwise hang a synchronous --check with no message; past the cap the child is killed and the kill becomes one more named skip.

The probe runs under the adapter's own injected HOME and environment, not the ambient process environment, so a relocated CODEX_HOME cannot verify ok against the operator's real ~/.codex.

A refuting probe means two different things, and which one it means depends on whether an artifact backs the registration:

The host answered "not registered", and Verdict Repair
this build wrote an artifact and that artifact still matches the canonical payload mcp-unreachable restart the runtime so it reloads its MCP configuration; re-applying would rewrite bytes that are already correct
the host's own registry IS the record (copilot mcp add leaves no file; a Codex config.toml that declares no tables has no artifact to be right) drift o2b install --target <t> --apply

A probe that could not RUN refutes nothing. Demoting a correct install because a binary was absent would be the same over-claim as a blanket ok, pointed the other way.

Tool ceilings, --host-target, and the tool-profile ladder (since v1.50.0)

The ceiling vocabulary has three states and unknown is never read as unbounded. declared carries a number and the place it is published; unbounded carries the place the ABSENCE of a limit is published; unknown carries the reason nobody has established one. Collapsing the third into the second is the misleading default the whole substrate exists to remove, because the layer above it selects a tool surface fail-open. Today exactly one row is declared - Cursor, at 40 tools across every enabled MCP server - and no row is unbounded; the member exists so a host that documents "no limit" is not recorded as unchecked.

o2b mcp --host-target <runtime> tells a running server which runtime launched it. Install adapters write it into the registration they generate; an unrecognised value is refused with exit 2 naming the known ids, rather than dropped - silently ignoring it would report the ceiling as unchecked on a host that publishes one. Its only other use is as the fallback for --harness below.

o2b mcp --harness <id> (since v1.70.0) names the harness the server runs under, so brain_context can render the operator's Brain/standing-rules/harness/<id>.md (see "Scoped operator rules" in how-it-works.md). The list is closed: aider, claude-code, codex, copilot-cli, cursor, gemini-cli, generic, grok, hermes, kiro, openclaw, opencode, pi. Without it the server takes --host-target, and with neither it matches no harness file. An unrecognised value exits 2:

$ o2b mcp --harness nope
o2b mcp: invalid --harness value: "nope"; expected one of: aider, claude-code, codex, copilot-cli, cursor, gemini-cli, generic, grok, hermes, kiro, openclaw, opencode, pi

Both refusals echo the value as a JSON string with every control character escaped; a refused --host-target value was echoed raw before v1.70.0.

The Claude Code plugin registers both of its servers with --harness claude-code, and the Hermes plugin launches its bridge with --harness hermes.

second_brain_capabilities reads it back and reports a host_ceiling object beside the tool counts: target, kind, max_tools, source, reason, and within_ceiling (the ADVERTISED count against the limit, which is what tools/list hands the host). A server nobody told - started by hand, or by a payload written before the flag existed - reports kind: "unknown" with a reason naming the flag and the command that regenerates the registration. An unchecked ceiling is never an absent one.

The tool-profile ladder has three tiers, highest first, and the generated registration is what it parameterises:

Tier Source
1 install.tool_profile in <vault>/Brain/_brain.yaml - the COMMITTED tier
2 mcp_tool_profile in the machine-local config.yaml
3 the host's own RUNTIME_FACTS row - catalog for Cursor, nothing for every other target

Tier 1 above tier 2 is the reverse of the guess, and it is the point. Install verification works by RE-CONSTRUCTION rather than a stored hash, so two machines cloning one vault must rebuild the same bytes; a stale key in one operator's ~/.config/open-second-brain/config.yaml outranking the file their teammate committed would make each machine report the other's correct install as drift.

There is no environment tier, and its removal is the reason to state the rule. OPEN_SECOND_BRAIN_MCP_TOOL_PROFILE used to sit above all three. It cannot: verify() rebuilds the expected artifact from the InstallEnv, so an input present on the apply invocation and absent on the next one makes a CORRECT install report drift - and the fix it offers would silently downgrade the profile the operator asked for. Every surviving tier is a file, and a file is still there tomorrow. The variable still selects the surface of the RUNNING server, which is the per-process act it is good for; it does not reach a generated registration.

The bottom tier resolving to nothing is not the same answer as full: it leaves the surface unnamed, which is the flag-free registration every host had before the ladder existed. An unknown profile NAME in either file tier is a hard refusal listing the valid names - the opposite of the running server, which fails open to the full surface rather than locking an agent out mid-session. Nothing is locked out here; an artifact is being generated, and generating one for a profile that does not exist installs a surface the operator never named.

Three of the ten targets carry neither dimension, because they write no MCP command line for a flag to land on: generic prints a payload for a host it was never told the name of, aider is wired through a managed YAML block and a sidecar context file, and pi is a skill symlink.

The committed install: block (since v1.50.0)

An optional install: block in <vault>/Brain/_brain.yaml carries the two settings that parameterise GENERATED output:

install:
  tool_profile: catalog        # full | writer | catalog | recall | minimal
  hook_timeout_seconds: 10     # 1..600; the generated Grok hook entry's cap

Both live in the vault rather than in the machine-local config for the reason the ladder above gives: verification is re-construction, so a knob only one machine can see makes the other report drift it cannot explain and re-apply a file that was already correct. hook_timeout_seconds bounds a generated lifecycle hook entry: 0 is not "no timeout" in any host that reads these entries - it is a hook killed before it can run - and ten minutes is the ceiling, past which the value is a unit error rather than an intent. Both bounds are hard errors at load, never clamped, because a clamped timeout is indistinguishable from one nobody set. Each key has a machine-local counterpart one tier down - mcp_tool_profile and install_hook_timeout_seconds in config.yaml - which the block outranks.

An unreadable _brain.yaml REFUSES rather than defaulting. An absent config is a vault with no settings to contradict; a malformed one is settings that exist and are not in force, and a writer cannot infer intent from bytes it could not parse. With the environment tier gone the vault tier is read on every resolution, so the refusal is unconditional - before, an environment variable short-circuited the ladder above it and a broken vault config went unnoticed for exactly those runs. o2b install --friction is the one surface that reports the refusal in a cell instead of raising, because it is a report and generates nothing.

Brain (observing memory)

o2b brain init                Bootstrap Brain/{inbox,preferences,retired,log,.snapshots}/ + _brain.yaml + _BRAIN.md; --starter drops the bundled example set
o2b brain feedback            Record one taste signal (--topic, --signal, --principle, --scope, ...); --scope is optional and falls back to feedback.default_scope from _brain.yaml when set
o2b brain dream               Run the deterministic consolidation pass (idempotent; usually cron'd - `o2b brain maintenance run --cron-template` prints a schedule that runs it inside the maintenance lane)
o2b brain apply-evidence      Record applied / violated against a preference for a durable artifact
o2b brain note <text>         Append a one-line narrative milestone to Brain/log/<today>.md (cron / shell mirror of brain_note)
o2b brain digest              Render a Markdown or JSON summary of recent Brain transitions; --window 7d for arbitrary lookback; Markdown links follow link_output_format / OBSIDIAN_LINK_FORMAT
o2b brain intent-review       Read-only pre-dream review of active signal clusters; --now ISO; --json mirrors brain_intent_review
o2b brain retention           Recommendation-only lifecycle review over retired preferences and processed signals; --now ISO; --json mirrors brain_retention
o2b brain monthly             Month-level Brain synthesis over timeline events, transitions, retirements, contradictions, and neglected areas; --month YYYY-MM; --json mirrors brain_brief view=monthly
o2b brain query               Read helper: by preference, by topic, or by log timestamp; --topic adds point-in-time recall over the expiration filter: --at <ISO|YYYY-MM-DD> evaluates expiration_date as of that instant (default now, unparseable is refused with exit 2, never coerced) and --show-expired keeps lapsed memories. Both apply to --topic only and are refused elsewhere; omitting both leaves the output byte-identical
o2b brain agent-query         Read source-agent provenance; filters by --agent, --topic, --query, --kind, --limit; --json mirrors brain_agent_query
o2b brain agent-diff          Compare source-agent coverage in browse/search/diff/map modes; --json mirrors brain_agent_diff
o2b brain writes              Recorded note writes, newest first: timestamp, operation, agent, device, target, and the content digest on each side; filters by --agent, --device, --path, --since, --until, --op; --json mirrors brain_writes action=list. The device comes off the log shard, so a legacy un-sharded log prints '-' rather than a guess
o2b brain writes revert       (CLI-only) Undo the recorded note writes a selector names: `(--agent A | --device D | --path P)` plus optional `--since` / `--until`, and `--apply <digest>` to run it. A selector naming none of agent, device or path is refused by name (`unbounded_selector`) - a time window alone is a vault rollback, which is `o2b brain rollback`. Without `--apply` it PLANS: each target resolves to `restore` (put the before-image of the oldest selected write back), `delete` (the selected writes created the note and nothing wrote it before) or a refusal naming `drift` (the bytes on disk are not the ones the newest selected write left), `interleaved` (an unselected write sits between the oldest and newest selected write in the merged `(timestamp, shard id, line)` order; line order inside one shard is append order, so a same-second write on the SAME shard is ordered rather than refused, while a same-second write from another device shard is refused on doubt), `image-missing` (the before-image is not in the store), `unrecorded` (the current bytes cannot be read at all) or `already-reverted` (the target already holds what the revert would produce). The plan prints its `digest:` on its own line followed by the exact `--apply` invocation, and exits 0 even when every target is refused - a plan that refuses everything is a report. `--apply` re-plans and refuses `digest_mismatch` before any byte moves, refuses `nothing_to_apply` when no entry is actionable, then runs every restore and delete inside one `note-revert` snapshot and records each as a note write of operation `revert`, so a revert is itself attributable and revertible. A frozen vault refuses it by name before anything is read
o2b brain writes prune-images (CLI-only) Remove note-write before-images older than --older-than-days (default 30); --dry-run lists what would go and removes nothing; --json
o2b brain reject              (CLI-only) Retire a preference; requires --reason "<text>". Subsequent signals on the same topic are suppressed.
o2b brain merge               (CLI-only) Fold one confirmed/quarantine pref into another (<keep> <drop>); --dry-run / --force; drop retires with reason 'merged-into'
o2b brain pin / unpin         (CLI-only) Toggle pinned: true on a preference (exempt from auto-retire)
o2b brain freeze              (CLI-only) Write Brain/.state/frozen.json and stop every content write in this vault (--reason "<text>", --json). Syncthing carries the marker, so the stop holds on every device that shares the vault, not only this one. The Brain log keeps recording: `freeze` names who set it, and each refused MCP call appends `write-refused`. Idempotent - a second freeze reports the first one's timestamp, agent and reason rather than replacing them. A marker that is present and unparseable freezes too: an operator who wrote a broken one meant to stop writes
o2b brain unfreeze            (CLI-only) Remove the freeze marker and reopen the content lane (--json). The `unfreeze` log event records who lifted it plus what the marker said, which is the only place those survive the file. Idempotent. `o2b brain rollback` is refused while frozen like every other content write, with no exemption for recovery: an operator restoring a frozen vault is the operator who should decide the freeze is over
o2b brain log verify          Walk every JSONL shard of the Brain log and report where each one stops linking up (--json). Each appended row carries `prev` and `h`: `h = sha256(canonical JSON of { prev, ts, kind, payload })`, `prev` the previous chained line's `h` or null at the shard's genesis, so an edited or deleted line is detectable. Output names the shard path, the line number, and the break - `hash-mismatch` (the line was edited), `prev-mismatch` (a line was removed or reordered), `malformed` (a line carries no chain link where the chain had already started). Rows written before the chain shipped carry no `h` and are counted as legacy: clean while they precede the chain, a break after it. The markdown twin is a derived rendering and is not chained. Reports only - nothing rewrites a log to make it verify and the read path never consults the chain - and exits 1 when any shard does not link up
o2b brain set-primary         (CLI-only) Declare or clear primary_agent in Brain/_brain.yaml (--clear)
o2b brain protect             (CLI-only) Emit / apply native deny rules for Brain/ (--target {claudecode|codex} [--apply])
o2b brain unprotect           (CLI-only) Remove the Open-Second-Brain-managed deny rules for the chosen target
o2b brain permissions         (CLI-only) Show the vault's trust policy document (Brain/_permissions.yaml) and a dry-run decision table resolving every declared agent against write/ingest/owner_write (`show`, --json adds the resolved rows), or list the decision ledger rows the gates append (`ledger --actor <name> --action <write|ingest|owner_write> --verdict <allow|ask|deny> --since <iso> --until <iso> --limit <n>`, --json). With no document every write is ungated; a document that cannot be read fails closed - `show` names the field and the file, and `o2b brain doctor` reports the same fault as `permissions-unreadable` with this verb as the exit
o2b brain snapshot log        (CLI-only) Newest-first listing of every recovery point: run id, created_at, typed reason, size, manifest presence, derived-store coverage; --reason filters (unregistered value exits 2), --limit caps, --json
o2b brain snapshot diff       (CLI-only) Read-only diff between two snapshots, or snapshot vs live Brain/
o2b brain rollback            (CLI-only) Restore Brain/ from a snapshot (--dry-run previews; drift abort vs --force-rollback); --list, the prompt and --json name the snapshot reason ('unknown' when the sidecar records none)
o2b brain upgrade             (CLI-only) Migrate release-owned files forward (_brain.yaml, _BRAIN.md, _OPEN_SECOND_BRAIN.md); --dry-run / --check / --apply --yes; --dry-run and --check also print the last failed automatic upgrade
o2b brain export              Read-only dump of active preferences, or (since v1.50.0) a session-transcript dataset: --format json|llms-txt|transcripts-jsonl [--out <path>] [--force]; the transcript form takes --transcripts <file|dir> and reads no vault at all (see "The transcript corpus" below). Since v1.49.0 the preference bytes pass the shared egress redactor (stderr carries a notice when anything was removed), and a `pref-*.md` the parser cannot read is REFUSED, not skipped: exit 1 naming every unreadable file in one run rather than exit 0 over a shorter list. `o2b brain doctor` reports the same files
o2b brain bank-export         (CLI-only) One-file backup bundle: preferences, the page graph, page contracts, the sources dashboard. Redacted on the way out; refuses with exit 1 on an unreadable `pref-*.md` (since v1.49.0)
o2b brain bank-import         (CLI-only) Restore a bank bundle (--mode skip|overwrite|merge) [--trusted-restore]. Preference rows restore UNTRUSTED by default: every row lands `unconfirmed`, unpinned, at low confidence, on a fresh trial window dated from the restore - including a row the bundle already marked unconfirmed - and the run names each row it reset; `--trusted-restore` keeps the carried status, confidence, pin and window verbatim, for a backup you vouch for. A malformed preference already in the DESTINATION does not abort the import: the rows restore and the run prints `topic-key check incomplete: <path>` for each rule the topic-collision scan could not read, because that list is then a partial answer (since v1.49.0)
o2b brain explorer            (CLI-only) Force-directed HTML graph of Brain/preferences + retired; live HTTP on 127.0.0.1 or --export <path> single-file. Keyboard-accessible listbox + localStorage layout persistence. Double-click a node to open it in Obsidian (live mode). Since v1.49.0 BOTH modes refuse a Brain file they cannot parse - `--export` writes nothing and exits 1, and live mode does not start, because a browser showing a silently smaller graph is the same lie as a written file - and `--export` runs the shared redactor over the graph, so its output differs from a pre-v1.49.0 export on any vault holding a credential-shaped value
o2b brain doctor              Check Brain-specific invariants (status-vs-folder, broken wikilinks, ...). --remediate [--dry-run] plans a dependency-ordered repair and applies auto-safe content-hash re-stamps
o2b brain health              Semantic-health report (since v0.14.0): contradictory confirmed preferences, recurring concepts with no dedicated preference, stale claims, plus a clean | watch | investigate verdict
o2b brain history             Render a preference's edit-history timeline (since v0.14.0): one entry per content mutation (principle / scope / status before -> after)
o2b brain audit               Render a preference's full mutation audit trail (since v0.21.0): create / promote / update / retire / merge with agent, reason, revision + content-hash before/after. ret- or bare-slug arg resolves to the same trail
o2b brain morning-brief       Read-only session-start summary (since v0.21.0): top confirmed preferences, recent reconcile open questions, recent notes; bounded by --max-chars-per-memory / --max-total-chars; --top-k / --lookback-days
o2b brain codec               Deterministic lossless session codec (since v0.22.0): --compress | --expand over stdin or --in <file>; structured content preserved byte-for-byte
o2b brain sources             Read-only dashboard of signals by (agent, source_type) (since v0.22.0): active/processed + distinct-topic counts; --json
o2b brain schema              Runtime schema report/admin (since v0.26.0): report|stats|lint|graph|explain|orphans|apply|sync; mutation writes are locked and audited
                              report names the pack's integrity: `ok` (digests to what the newest audited apply recorded), `modified` (with both digests), or `unverified` with the reason - `no-apply-recorded`, `audit-unreadable`, or `config-absent`. The digest is over the rendered `schema:` block, so comment and whitespace churn elsewhere in `_brain.yaml` never moves it
                              sync is NOT implemented and exits non-zero. It previously returned `{updated: 0, skipped: 0}` with a "no backfill was required" note without reading the vault at all, so a caller could not tell "nothing needed doing" from "nothing was done"
o2b brain schema apply        --mutation '<json>' (repeatable) applies audited, locked mutations to Brain/_brain.yaml; --actor NAME --reason TEXT annotate the audit record
                              --dry-run previews instead (since v1.43.0): it returns the pack that would result plus its leaf-level diff and writes nothing at all - no config write, no audit record, no lock file. Before v1.43.0 `apply` parsed the flag and mutated anyway. The preview runs the same pack validator the apply runs, but NOT the two checks that exist only because the apply writes (the vault-identity write guard, and the atomic writer's re-parse of the rendered YAML), so a batch that renders to unparseable YAML previews clean and fails on apply
o2b brain watchdog            Probe Brain config/dirs/search index and plan safe recovery (since v0.26.0): --remediate [--dry-run], --restore <run_id> [--force-restore], --json
o2b brain graph-export        Serialise the vault knowledge graph (pages, wikilinks, typed relations) to a stable graph.json (since v0.22.0): stdout or --out <file>
o2b brain graph-import        Reconstruct vault page stubs from a graph.json (since v0.22.0): --mode skip|overwrite|merge; vault-guarded writes
o2b brain backlinks           List inbound references to a Brain artifact id
o2b brain semantics-backfill  Dry-run typed preference-edge backfill preview (since v0.24.0): --json returns missing inverse superseded_by proposals; no writes
o2b brain mcp-landscape       List MCP servers configured across the vault (since v0.19.0): name, source file, packages, required env-var names (values never read)
o2b brain scan-inline         Capture `@osb` markers from folders listed under `notes.read_paths` in _brain.yaml
o2b brain today               Read-only today dashboard: due and overdue obligations, open loops (`@osb loop <text>` markers, closed by `@osb loop close id=<id>`), merged recent activity, totals; --lookback-days (default 7), --limit (default 20), --json. Each section derives live; a failing section is reported and the rest still render
o2b brain apply-markers       Turn `@osb set note=<target> field=<field> value=<value>` markers into schema-validated frontmatter writes. Report mode by default (writes nothing); --apply needs `guardrails.marker_writeback` in _brain.yaml, consumes applied markers so a re-run is idempotent, and leaves unresolvable or refused targets unconsumed; --path <subdir> (repeatable) or notes.read_paths
o2b brain import-session      Replay signals from a registered agent session .jsonl (or directory); --recall also stores turns in the session recall DAG. Since v1.50.0 the path argument has a machine-wide sibling: --status reports per-runtime coverage, --discover lists what has never been imported, --discover --all imports it (see "Session logs this machine already has" below)
o2b brain session-hook        Internal hook bridge: read one lifecycle payload from stdin, capture prompt markers / brain_feedback, append lifecycle audit/log rows
o2b brain context-receipts    List/show opt-in prompt context receipt continuity records (since v0.29.0)
o2b brain recall-telemetry    List/summarise opt-in recall telemetry continuity records (since v0.29.0)
o2b brain generation-reports  Record/list/summarise opt-in inbound LLM generation traces (prompt hash + token counts only; kernel never calls an LLM)
o2b brain context-presets     Show/suggest/diff read-only context budget presets (since v0.29.0)
o2b brain pre-compact-extract Extract decision/commitment/outcome/rule/open-question continuity records from bounded text (since v0.29.0)
o2b brain post-compact-audit  Audit pinned-anchor survival after a compaction and re-assert drifted anchors (gated by post_compact_survival_audit)
o2b brain session-grep        Search imported session recall raw turns and summary nodes (since v0.29.0)
o2b brain session-describe    Count raw turns and deterministic summary depths for one session recall DAG (since v0.29.0)
o2b brain session-expand      Expand a session recall node to immediate sources and paginated raw turn content (since v0.29.0)
o2b brain session-summary     write|get|list a session-scoped structured digest (request/decisions/learnings/next_steps) (since v1.11.0)
o2b brain idea-lineage        Trace how a derived artifact was reached: observation -> synthesis -> conclusion over continuity sourceRefs or belief-evolution (since v1.11.0)
o2b brain note-lifecycle      <rename|move|archive|delete> <path> [<to>] [--apply] [--confirm] [--expect N] [--strict] [--json] - note-FILE lifecycle; source and destination both go through the create-note path envelope. Dry-run without --apply; delete also needs --confirm. A relocation rewrites inbound [[wikilinks]] vault-wide (both path spellings always, the bare basename only when exactly one note carries it) and the receipt names how stale the search index still is; delete rewrites nothing, reports the references it strands, and returns a recoverability verdict saying the archive it took does not cover a note outside Brain/
o2b brain scaffold-stub       list [--limit N] [--json] | write <target> [--path <p>] [--source <s>...] [--if-exists refuse|skip] [--apply] [--json] - unresolved wikilink targets. list reads them from the search index and refuses with a state and a next_command when that index is missing or only partially resolved, never printing an empty list for an index nobody finished resolving; write materialises a stub whose title is the target and whose body links back to its sources. A target that already resolves, or names several notes, is refused
o2b brain note-history        Decompose a note's git history into episodic phases split on a commit-time gap (since v1.11.0)
o2b brain import-claude-memory (CLI-only) Import metadata.type=feedback entries from a Claude Code memory directory into Brain/preferences/. --dry-run / --apply, sidecar manifest for idempotency, UPDATE preserves accumulated evidence; --approval-digest <hex> seals the plan - non-interactive --apply must carry the digest a prior --dry-run printed and then lands exactly the approved plan or nothing (a stale digest refuses with a re-run remedy), an entry whose name cannot land a parseable tag skips as a named row instead of aborting, and a digest paired with --dry-run refuses

Time axis (since v0.10.18)

o2b brain timeline            Chronological event list; filter by --pref-id / --topic / --kind / --since / --until / --limit
o2b brain evolution           Per-preference or per-topic belief evolution with running evidence counts, walks supersedes / superseded_by retire chains
o2b brain stale               Structural staleness report (preferences / signals / log files) using configurable per-kind thresholds
o2b brain daily               Daily brief: counters, transitions, source pointers
o2b brain weekly              7-day synthesis with contradictions list (signal-suppressed events + apply-evidence violated rows)
o2b brain monthly             Monthly synthesis: event count, status transitions, retirements, contradictions, neglected areas

Maintenance and operator surfaces

o2b brain actions             Ranked next-step list combining doctor / dream warnings and lint candidates
o2b brain summary             Operator dashboard - trust verdict (clean | watch | investigate), doctor / dream counts, verification delta, instruction-file ceiling warnings, top maintenance actions
o2b brain page-dedup          Page-level duplicate detector (by content hash + frontmatter similarity)
o2b brain doctor              gains the `merge-chain-dangling` warning (since v1.64.0): a `merged_into:` pointer whose canonical resolves to no file - the page silently dropped out of dedup candidacy forever; chains crossing Brain/retired/ are walked, and cycle, depth and malformed intermediates stay with `o2b brain lint --consolidate`'s unresolved list; the repair is a content judgement (re-point at a surviving canonical, remove the pointer to un-merge, or restore the file), so the finding names the judgement instead of a next command
o2b brain lint                Self-healing structural drift fixer; --consolidate folds multi-source duplicates
o2b brain token-footprint     Token-budget monitor across instruction files and active.md
o2b brain context-pack        Bounded-token vault slice for priming an agent's context window (--max-tokens N, --lanes for directives/constraints/consider; v0.29.0 adds opt-in --receipt, --telemetry, --cache-stable, --dedup-repeated)
o2b brain synthesise          Concept-scoped JSON envelope: target node + linkers + optional unlinked mentions
o2b brain moc-audit           Per-MOC coverage audit: classify cluster members into well-covered / fragile / candidate-missing
o2b brain unlinked            Raw-text mentions outside `[[...]]` (Unicode-aware boundaries)
o2b brain skill-proposals recover [--discard-unreadable]
                              Resolve accept sequences a crash abandoned, rolling each back to its pending draft or forward to rebuilt projections. Refuses by name, naming the exact file, on a held `Brain/skill-proposals/accept.lock` (nothing breaks a lock - a live writer cannot be told from a crashed one) and on an accept-journal marker that cannot be parsed; `--discard-unreadable` removes those markers, which unblocks accepting without claiming the sequence each marked was resolved

Recovery is a vault-wide sweep, not a per-slug one: recover - and every accept, which runs the same sweep first - resolves EVERY outstanding journal it finds, so accepting one proposal can roll another slug's abandoned sequence back. That is deliberate. The sweep and the accept share one vault-wide lock, which is what makes it safe, and leaving a crashed sequence unresolved would let the next accept build on a duplicate. It writes only under Brain/skill-proposals/ and Brain/procedures/, and only files the abandoned sequence itself created - the journal records whether each target pre-existed.

Context continuity and receipts (since v0.29.0)

o2b brain context-pack        Existing budgeted context pack; add --receipt to emit a context receipt, --telemetry to emit redacted recall telemetry, --cache-stable for stable ordering diagnostics, and --dedup-repeated for repeated-context reference hints
o2b brain context-receipts    list [--trigger context_pack|pre_compress] [--host <name>] [--session-id <id>] [--limit <n>] [--json]; show <receipt-id> [--json]
o2b brain recall-telemetry    list|summary [--mode search|context_pack|pre_compress|query] [--status ok|empty|error|timeout] [--channel mcp|cli|hook] [--host <name>] [--since <iso>] [--until <iso>] [--limit <n>] [--json] - --channel is a read filter on list and summary only; gate-list, gate-summary and cost refuse it by name rather than ignoring it
o2b brain generation-reports  record <write_session|context_pack|dream_stage> --ref <id> --agent <name> --prompt <text> [--enable] [--provider <p>] [--model <m>] [--finish-reason <r>] [--latency-ms <n>] [--input-tokens <n>] [--output-tokens <n>] [--cached-tokens <n>] [--total-tokens <n>] [--scope <s>] [--source <id[=path]>...] [--created-at <iso>] [--json]; list|summary [--handoff <kind>] [--agent <name>] [--since <iso>] [--until <iso>] [--limit <n>] [--json]; show <report-id> [--json] - record is gated (default off) by --enable or generation_trace_enabled; stores prompt_hash + counts only
o2b brain context-presets     show [tight-context|long-context] --json; suggest --model <name> --context-window <tokens> --json; diff <preset-id> [current-value flags] [--override <path>...] --json
o2b brain pre-compact-extract --session-id <id> --turn-start <id> --turn-end <id> --text <bounded-text> [--host <name>] [--max-chars <n>] [--json]
o2b brain post-compact-audit  [--session-id <id>] [--no-reassert] [--force] [--vault <path>] [--json] - reads { session_id, messages } JSON from stdin; gated off by default (post_compact_survival_audit), --force overrides
o2b brain import-session      <path> --recall [--recall-session-id <id>] [--recall-summary-group-size <n>] [--json]
o2b brain session-grep        --query <text> [--session-id <id>] [--limit <n>] [--snippet-chars <n>] [--json]
o2b brain session-describe    --session-id <id> [--json]
o2b brain session-expand      <record-id> [--raw-limit <n>] [--cursor <offset>] [--json]
o2b brain handoff             <session-file> [--session-id <id>] [--format auto|claude|codex|hermes] [--json] - write Brain/handoffs/<date>-<scope>.md (since v0.37.0)
o2b brain hygiene             scan | apply --ids <id,...> [--detectors conflicts,dedup,freshness,usefulness,slug-collisions,tags,capture-scope] [--dry-run] [--json] - hygiene findings pipeline; review findings never execute (since v1.3.0); the default sweep runs every detector except the opt-in ones (since v1.64.0: `slug-collisions` is default-on and reports same-stem groups such as topic.md beside topic-2.md as info findings, `tags` is opt-in and audits inline body tags only - frontmatter tags arrays are out of scope; since v1.67.0 `capture-scope` is default-on and warns, with proposed action review, when a retrievable page of active knowledge cites only sources that are currently url-only)
o2b brain hygiene             scan: with the optional `dedup` decision-model use in enforce, dedup findings carry an advisory `decision_model` verdict and a confident `different` is listed last; `o2b brain doctor` annotates `entity-alias-candidate` warnings the same way; `apply` ignores verdicts
o2b brain refresh             --stale [--dry-run] [--json] - targeted recompile of stale derived pages; orphans archive into Brain/.snapshots (since v1.3.0)
o2b brain anticipate          --session <id> [--refresh] [--signal <text>] [--json] - read or warm the anticipatory context cache for the session's lineage root (since v1.3.0)
o2b brain intention           set|show|list|move [--scope S] [--text T] [--json] - scoped current-intention chains under Brain/intentions/ (since v0.37.0)
o2b brain obligation          add|done|list|show|remove [--title T] [--cadence C] [--anchor YYYY-MM-DD] [--date YYYY-MM-DD] [--slug S] [--notes N] [--overdue] [--json] - recurring obligations under Brain/obligations/ with a deterministic cadence-driven next-due date; cadences: daily|weekly|biweekly|monthly|quarterly|yearly|every-<N>-days (since v1.15.0)
o2b brain agenda              --events <file|-> [--focus-min N] [--owner-domain D[,D2]] [--workday-start HH:MM --workday-end HH:MM] [--json] - stateless agenda synthesis over caller-provided calendar events: overlap conflicts, free focus blocks, external-organizer flags; no vault writes (since v1.15.0)
o2b brain okf-export          --out <dir> [--force] [--json] - write a portable Open Knowledge Format bundle (concepts/queries/references + date-grouped log.md + okf.json manifest) to a directory; read-only on the vault (since v1.15.0)
o2b brain okf-import          <bundle-dir> [--trusted] [--json] - import an Open Knowledge Format bundle; default stages pages under OKF Review/ as review candidates, --trusted writes them to their recorded paths; foreign-producer provenance is stamped (since v1.15.0); the manifest producer is never trusted - without --trusted, machinery frontmatter (_status, owner, origin_channel, ...) is stripped from every bundle, and only --trusted admits the Brain/sources/ and Brain/reports/ lanes

Receipts, telemetry, transforms, and session recall import are opt-in. Receipt and telemetry records store redaction-safe payloads, source references, hashes, counters, and bounded snippets rather than raw private prompt context; session recall stores redacted turn text only when explicitly imported for later expansion.

Workspace reach and proactive insight (since v0.38.0)

o2b brain project             link <path> [--vault V] | list | remove <path> | status [path] [--json] - .o2b-vault.json pointer files; resolveVault honours the nearest pointer (env override still wins)
o2b brain source              add <vault> --alias <name> | list | remove <alias> [--json] - read-only recall sources of the active vault; BROKEN flag for missing targets
o2b brain links               normalize [path-prefix] [--mode preserve|full|short] [--write] [--json] - wikilink path format rewrite; dry-run by default; config key wiki_link_format
o2b brain profile             [--stale-seconds N] [--force] [--json] - materialize Brain/profile.md digest + .o2bfs root marker (age-gated)
o2b brain sgrep               <query> [path-prefix] [--limit N] [--keyword-only] [--json] - grep-shaped semantic search; path:line: lines; exit 1 on no matches
o2b brain trigger             scan | list [--status S] | ack <id> | dismiss <id> | act <id> | suppress <id> | unsuppress <id> | history [--json] - grounded trigger queue; cooldown via trigger_cooldown_days (default 7); suppress silences a finding indefinitely and unsuppress restores the status it interrupted; list reports the suppressed count; scan, list and history all print `unreadable: <n>` followed by one line per record the store could not parse - printed at zero too, so an omitted line can never read as a clean queue, and printed BEFORE "no open triggers" so that line is never the only thing said about a queue holding something unparseable
o2b brain deep-synthesis      <topic> [--limit N] [--triggers] [--json] - deterministic topic dossier: agreements, contradictions, stale claims, knowledge gaps, plus a strongest-objection steelman
o2b brain ideas               [--cap N] [--triggers] [--json] - ranked next-direction candidates from open questions, orphan notes, aging signals
o2b brain recall-telemetry    gate-list | gate-summary [--host <name>] [--since <iso>] [--until <iso>] [--limit <n>] [--json] - recall-gate decision telemetry (recall_gate_telemetry, default off)
o2b search <query> --global   Cross-vault union over profiles + read-only sources; origin-labelled results; external vaults are never written to

Trigger generation, brief delivery, and gate telemetry are pull-based or config-gated; pointer resolution activates only when a pointer file exists.

Memory observability (since v0.39.0)

o2b brain continuity          export --format atof|atif [--session <id>] [--month YYYY-MM] [--out <dir>] [--json] - read-only trajectory export of the continuity store; private records dropped, redacted text stays masked
o2b brain bench               memory --fixture <name|path> [--resume <run-id>] [--runs-dir <dir>] [--json] - memory quality benchmark over a disposable fixture vault; quality, latency, and context cost reported separately; checkpoint/resume by run id; exit 1 on any failed question

The benchmark never touches the configured vault - the fixture materializes into <runs-dir>/<run-id>/vault (default runs dir .open-second-brain/bench-runs/, gitignored). The optional bench_judge_cmd config key (env OPEN_SECOND_BRAIN_BENCH_JUDGE_CMD) arms an advisory external judge; absent means the judge phase is skipped. The full observability contract (event kinds, gates, correlation ids, payload safety, schema version) lives in docs/observability.md.

Externalized session payloads (payload registry)

o2b brain payload             get <osb-payload://sha256> [--offset N] [--limit N] [--json] - one page of the exact stored content (default 4000 chars); raw to stdout without --json, next offset on stderr
                              list [--json] - stored payloads with byte size and reference count (orphans flagged), plus referenced payloads whose file is gone
                              gc [--apply] [--json] - dry-run by default; --apply removes only payloads nothing in the vault references and older than ten minutes (younger ones are listed as `deferred`), behind a payload-gc recovery point and under the payload store lock session import also takes

import-session --recall bounds every recalled turn before its continuity row is written. A data URI or base64 run longer than sessions.payload_max_inline_chars (default 512) moves to Brain/.payloads/<sha256>.txt; turn text still longer than sessions.payload_max_text_chars (default 32000) moves out whole - which is what catches a giant plain tool output - and the row keeps a 1000-char head preview. Either way the row holds [payload: osb-payload://<sha256> chars=N] and lists its refs in payload_refs, so recall search and summary nodes only ever see the placeholder. A turn under both bounds is stored exactly as before.

Payload bytes are private-region-stripped and redacted (redactRawOutput, no scan-window cap) before they are written, so a page reads back the stored redacted text byte for byte. Brain/.payloads/ is refused by the index-admission predicate, is archived by every snapshot (a restore keeps restored refs resolvable), and is not exported by bank-export or okf-export - an exported page keeps its placeholder. "Referenced" means named by any text file in the vault outside Brain/.snapshots/, .git, node_modules and .stversions, and transitively by a live payload (a whole-turn payload names the blobs moved out of it). o2b brain doctor reports payload-orphan, payload-missing and continuity-row-oversized (rows written before the registry existed).

Project history (since v0.40.0)

o2b brain git                 ingest <repo-path> [--max-count N] - read-only walk of a worktree; commit/tag/release records + digest note land under Brain/projects/git/<repo-key>/; incremental via SHA-validated watermark, full re-scan with a reported warning on force-push or tampered state
                              status - per-repo watermarks and record counts
                              find [text] [--repo K] [--file F] [--author A] [--since S] [--until U] [--limit N] - query ingested history newest-first; no live git on the query path
                              mine [--repo K] - surface decision-shaped commits as draft ADR candidate notes under Brain/decisions/candidates/ (sha-stable identity, skip-existing)
o2b brain architect           <project-path> [--progress] - deterministic project scan (built-in runtime only, no dependency, no LLM) rendered as architecture notes under Brain/projects/arch/<repo-key>/; generated content lives in o2b:begin/o2b:end sentinel regions, operator prose outside regions survives every re-scan byte-for-byte; dependency manifests (package.json, pyproject.toml, Cargo.toml, go.mod) are read at the root and at each module, pom.xml, build.gradle, Gemfile and composer.json are reported unsupported, each manifest is read, malformed, unreadable or unsupported; --json always carries manifests: [{path, ecosystem, status, detail?}] sorted by path, and text mode adds one "manifests not read" line only when some manifest was not read; the overview gains dependencies and module-dependencies regions, module notes a dependencies region and a generator-owned depends_on frontmatter key

All flags accept --vault V and --json. The ingest never modifies the scanned repository; every caller-supplied sha is validated against the full-40-hex grammar before it can reach a git argument.

Agent write sessions (since v0.41.0)

o2b brain session             open --target <Brain/...md> [--schema-type S] [--intent create|overwrite|merge] [--prompt P] [--require-review] [--retry-cap N] - open an artifact write session; the envelope carries the generation prompt, schema hints, and collision metadata for an occupied target
                              submit <id> [--file F|-] - submit the generated artifact (stdin without --file); done | needs-correction with coded errors and a compact correction prompt | needs-review
                              approve <id> - operator-side commit of a needs-review session
                              abandon <id> - terminal abandon
                              status <id> | list | sweep - inspect or clean the session store (Brain/.sessions/write/, lazy TTL default 24h)
o2b brain panel               open <topic...> [--personas a,b,c] [--target T] [--require-review] - convene a decision panel; personas from Brain/personas/ (built-in defaults: technical, strategic, risk, user-experience)
                              submit <id> [--file F|-] - answer the current persona step or the synthesis; the committed note lands under Brain/decisions/panels/
                              status <id> - live envelope of a panel session

Envelopes are stable JSON with --json (status, step, prompt, errors, attempts_left, expires_at, target_path, existing) - the same contract the MCP brain_write_session tool returns. create intent never overwrites an existing target; merge appends a session-stamped delimited section; a target must sit under Brain/sources/, Brain/reports/, Brain/distillations/, Brain/notes/ or Brain/decisions/panels/ (case-folded); the rest of Brain/ is refused with target-reserved, and the commit re-checks where the bytes land and applies the write binding. The Brain never generates content - the calling agent does.

Recall activation (since v0.42.0)

o2b brain activation          status [--top N] - folded activation state: event/path/co-access counts plus the strongest paths
                              sweep [--retention-days N] [--max-events N] - drop access events outside the retention window or beyond the newest-N cap and refold (--max-events 0 clears every retained event)

CLI and MCP searches record which documents they surfaced as one JSON event per access under Brain/search/activation/ (query hashed, never raw text). o2b search <query> --no-record-access suppresses recording for one query; the MCP brain_search tool accepts record_access: false. Cross-vault (--global) and query-cache-hit searches never record, so reinforcement is miss-driven. The derived Brain/search/activation-state.json is a replayable fold - deleting it loses nothing. search_activation_enabled: false disables both the boost and recording; search_two_pass_enabled: false disables the evidence-pack broadened retry.

Entity truth and self-improving dream (since v0.43.0)

o2b brain truth               ingest --entity E --aspect A --value V --source S [--quantity-value N --quantity-unit U --quantity-action W] - append one claim to the ledger
                              slots [--entity E] - current values with superseded history and CONTESTED flags
                              conflicts [--window-days N] - value conflicts (independent sources within the window; resolution always ask_user)
                              aggregate --action W [--unit U] [--entity E] - sum exact (entity, action, unit) quantity matches
                              collisions [--window-days N] - cross-agent convergence on one entity
                              sweep [--max-events N] - keep the newest N claim events and refold
o2b brain facts               decompose (--file <path> | --text <text>) [--ingest --entity E] - deterministic atomic assertions; --ingest appends structured-family claims
o2b brain dead-end            record --approach T --reason T [--context T] | list - negative-knowledge registry under Brain/dead-ends/
o2b brain foresight           [--horizon-days N] [--write] - forward projection: routines coming due, open commitments, open questions

Claims live as device-sharded append-only JSONL under Brain/truth/ with a recomputable state.json fold - deleting the cache loses nothing. The merge guard rides o2b brain merge (an entity-guard refusal when the two preferences anchor disjoint people/orgs; --force bypasses). o2b brain apply-evidence accepts --outcome success|failure|unknown; the dream pass stages outcome_regressions with a deterministic confidence penalty when applied events carry repeated failures. brain_review_candidates annotates inbox signals with signal_novelty when the vault has indexed embeddings.

Write-time integrity and governance (since v0.44.0)

o2b brain label               <path> <dimension>=<value> | --remove <dimension> | --show - controlled-vocabulary classification; fail-closed against the schema pack's labels field
o2b brain label               <path> --suggest [--dimensions a,b] - read-only advisory suggestions from the optional `labels` decision-model use; never assigns; `available: false` while the use is off; a private note is refused (see docs/decision-models/labels.md)
o2b brain attr                <path> <field>=<value> | --remove <field> | --show - per-type attribute fields; an undeclared field error lists the declared fields WITH descriptions
o2b brain tiers               check | restore <path> [--field F] --apply | accept <path> [--field F] - staged repair for identity-tier frontmatter hand-edits
o2b brain secret              set <name> [--env-var V] [--allow PATTERN]... [--from-env SRC] | list | rm <name> | run <name> -- <command...> | lock | unlock [--passphrase-from-env SRC] | export --out FILE [--passphrase-from-env SRC] | import FILE [--replace] [--passphrase-from-env SRC] [--vault <path>] - capability-gated custody; the value enters via stdin, never argv. `unlock` wraps the keyfile under a passphrase (read from stdin or --passphrase-from-env, never argv) and holds the key for this process only; `lock` clears that holder; `export` writes every entry as one passphrase-encrypted bundle to `--out` (values re-encrypted, names and env-var mappings travel in the clear past the shared egress redactor); `import` restores a bundle, refusing names the store already holds unless --replace
o2b brain maintenance         run [--force] [--retry <task|custom:name>] [--force-cost] [--window H-H] [--tz ZONE] [--busy-minutes N] [--busy-threshold N] [--progress] | status [--limit N] - quiet-window lease-guarded lane for dream, reindex, bridges, clusters and declared custom tasks
o2b brain maintenance         run --cron-template [--interval N] [--format cron|systemd] [--window H-H --tz ZONE] - print the cron or systemd recipe that schedules the lane (default interval 1h); prints and installs nothing

The schema pack gains four additive ontology fields (labels, link_constraints, attributes, frontmatter_tiers) with audited mutations through o2b brain schema apply. Link constraints enforce at index materialization: a typed edge whose endpoint page types violate the declared pairs falls back to an untyped link, o2b brain schema lint lists each violation, and removing the constraint restores the edges on the next index run. Tier drift detection rides the same index pass - the snapshot keeps the expected value, so reindexes never absorb a hand-edit, and brain_doctor warns with the open count. Filter labelled recall with o2b search <q> --property labels=<dim>/<value>. Secrets protect against context leakage and vault sync exposure, not against root; every custody operation lands a no-values record in Brain/log/secret-custody/. A maintenance gate skip exits 0 so cron never alarms on a quiet hour.

The lane's two vault-side knobs live in Brain/_brain.yaml under maintenance:, because they answer what the cron line cannot. host_pressure_percent adds a fourth gate: skip when the host's one-minute run queue stands at or above that percentage of the CPUs this process may use. Unset by default, which leaves the gate off. Where the metric is degenerate - a platform whose load average is a constant, or a cgroup with a CPU bandwidth quota, where the run queue is the whole host's - the gate stays open and the journal carries a separate pressure:unmeasurable row naming the reason, so an unreadable host is never reported as a quiet one. failure_streak_limit (default 3) refuses a lane task that has failed that many times in a row in the run journal, naming the streak; a single journaled success clears it. The refusal is per task - the other three still run under the same lease - and there are two ways past it: --retry <task> (repeatable) attempts just that task with every gate the operator configured still in force, and --force runs the whole lane past every soft gate and every refusal. Both escapes exist on the MCP surface too - brain_maintenance takes retry_tasks, busy_minutes, busy_threshold and a status limit, so an agent reading the refusal can act on it without reaching for force, which switches off three gates it never meant to touch. A refused task is reported as REFUSED, never FAILED, and the run exits 7 rather than 1: nothing was attempted, so nothing failed, but a standing refusal is not the quiet hour that exits 0 either. An attempted failure still exits 1 and outranks a refusal in the same run. The refused:streak journal row carries the count it refused on, so the streak does not silently reset when its evidence rolls off the journal cap.

Since v1.64.0 the lane's timeouts are honest and its embedding spend is named. A task killed at its safeguard deadline exits 6 (probeIncomplete), keyed on timed_out rows alone: the only proved fact is that the pass did not finish, which is not the same as a task failing; a proved failure (1) outranks a timeout, a timeout outranks a refusal (7), and the render says TIMED OUT with the safeguard detail. The reindex task stays keyword-only unless maintenance_embeddings: true (env OPEN_SECOND_BRAIN_MAINTENANCE_EMBEDDINGS, default false) opts it in; then it requests the embedding phase whenever the resolved semantic config can reach a provider, announces the pending-spend estimate before the pass, and journals the phase's own cost-gate result as the per-run receipt - model, tokens, estimated cost, and whether a bypass fired - on the task row, the journal line and the maintenance_spend metrics surface; the preview is an estimate by position, the receipt prices what the completed pass embedded, and a run killed mid-spend receipts nothing. --force-cost (MCP force_cost) bypasses a positive embedding_cost_gate_usd for this run, recorded on the receipt when it overrode a gate that would have refused. A safeguard timeout is journaled as timed_out and neither counts toward nor resets the failure streak that refuses a task.

Since v1.72.0 the spend is priced honestly. The banner prints price unknown instead of $0.0000 when the embedding model has no known price, the receipt and the maintenance_spend metric carry price_source (builtin, operator or unknown) with a null estimate for an unknown price, and o2b brain maintenance status lists each receipt with its price_source (unrecorded for rows journaled before this release). Under a positive embedding_cost_gate_usd an unpriced model refuses the reindex task's embedding phase with EMBEDDING_COST_UNPRICED unless --force-cost; declare the price with embedding_price_model and embedding_price_usd_per_mtok (see "Embedding prices" below).

Since v1.65.0 the lane prints its own schedule and can carry an install's own upkeep. o2b brain maintenance run --cron-template prints a script (~/.local/bin/osb-maintenance-<hash>.sh, where <hash> is the first 8 hex characters of the SHA-256 of the resolved vault path, so each vault gets its own script, Hermes job and systemd units and re-rendering for the same vault keeps the name) and its scheduler lines - a crontab line and the Hermes form by default, a systemd user timer with --format systemd - with --interval <N>m|h|d defaulting to 1h: the gates decide whether work happens and a gate skip exits 0, so an hourly schedule plus --window is the intended pattern. The script embeds the resolved vault, runs o2b brain maintenance run --vault '<vault>' --json, stays silent on exit 0 and prints the captured JSON and keeps the exit code otherwise. The rendered script calls o2b by name, so it expects o2b on the scheduler's PATH: cron jobs and systemd user services start with a minimal PATH, so add the install directory (usually ~/.local/bin) to that PATH. The verb returns before the lease, the gates, the journal and the metrics, so printing a recipe leaves no trace. A bad interval, window or format, --interval or --format without --cron-template, a lane-run flag (--force, --retry, --force-cost, --busy-minutes, --busy-threshold, --agent, --progress, --json) beside --cron-template, and status --cron-template exit 2.

Custom lane tasks (since v1.65.0) put an install's own upkeep under the lane's window, busy, pressure, lease and streak gates. They are declared in the machine config file, never in the vault: maintenance_custom_<name>: <command>, optionally maintenance_custom_<name>_cwd: <absolute dir> (default the running user's home directory; the vault is allowed when named) and maintenance_custom_<name>_timeout_seconds: <N> (default 120; a whole number from 1 to 1200). The lease is taken once per pass for 30 minutes, so the declared custom timeouts together stay within 1200 s, which leaves 600 s for the built-in tasks; in name order, a task whose timeout would take the sum past 1200 s is refused by name, and the full 8 tasks fit at the default. Declared tasks run only with the master switch maintenance_custom_tasks: true (env OPEN_SECOND_BRAIN_MAINTENANCE_CUSTOM_TASKS, default off; 0 turns it off on one host). Each runs as custom:<name>, where <name> matches ^[a-z][a-z0-9-]{0,31}$ (no underscores, so the suffixes stay unambiguous, and no task can be named tasks), after the four built-in tasks and stale-first with them; at most 8 are declared. The command runs through sh -c (cmd.exe /d /s /c on Windows) with stdin closed, stdout discarded and O2B_VAULT set. Its environment is this process's minus every variable whose name declares a credential (*_API_KEY, *_TOKEN, *_SECRET, *PASSWORD* and the other names the redactor treats as secrets); PATH, HOME, LANG, LC_*, TZ, TMPDIR and O2B_VAULT are always kept, and a command that needs a key reads it from its own configuration. The default working directory is not the vault because the vault syncs and is writable through the MCP write tools, and on Windows a bare command name resolves in the working directory first; use absolute executable paths. A _cwd that does not exist fails the task as custom task <name>: cwd does not exist: <path>. A non-zero exit is reported as exit <N>: <stderr tail>, redacted and capped at 4096 bytes. The task's outcome is its shell's exit status: it does not wait for processes the command left in the background, but those still belong to the task, and on POSIX the command's whole process group is killed (SIGTERM, then SIGKILL a second later) when the timeout elapses or the o2b process exits, whichever is first, also after the shell has exited (on Windows taskkill /T /F kills the tree while the shell runs). Past its timeout the run is journaled as timed out, and unlike a built-in task, a custom task's timeout counts toward its failure_streak_limit and exits 1 rather than 6, so a command that always hangs is refused like one that always fails. A shell that exits before its timeout keeps its exit status even when a background process still holds its stderr at the deadline. --retry custom:<name> (MCP retry_tasks) attempts a refused custom task alone; an undeclared name is refused with the registered list. A bad declaration (a bad name, an empty command, an out-of-range timeout, a relative _cwd, a suffix key without its command, a ninth task, a timeout over the budget) is named on stderr as custom task refused: <reason> (MCP custom_task_errors) and the valid ones still run; status says custom tasks declared but maintenance_custom_tasks is off while the switch is off, or custom tasks declared but OPEN_SECOND_BRAIN_MAINTENANCE_CUSTOM_TASKS turns them off when the env override did it. A custom task never calls a model on the lane's behalf and records no spend receipt, and the lane cannot police spend or side effects inside an operator's own command. No MCP parameter can add, edit or read a command. The config reader strips one pair of matching surrounding quotes from a value, so a command that starts and ends with the same quote character loses them; wrap such a command in single quotes (maintenance_custom_tidy: '"/opt/my tool" --flag "x"').

Link and recall intelligence (since v0.45.0)

o2b brain bridges             discover [--max N] [--min-similarity X] [--progress] | list | accept <source> <target> | dismiss <source> <target> - embedding-near link proposals over the vec index, reviewable artifact, accept writes one related: wikilink
o2b brain clusters            run [--min-size N] [--batch-size N] [--if-stale] [--progress] | list - graph-wide community detection; derived digests under Brain/clusters/, regenerated per run; --batch-size materializes in chunks with isolated, reported per-batch failures; --if-stale recomputes only when the freshness verdict is not `fresh` (see "Materialization freshness" below)
o2b brain vitals               [--orphan-threshold N] - aggregate governance scorecard over confirmed preferences: domain_diversity (scope entropy), connectivity_index (mean evidenced_by count), orphan_preferences (below threshold, default 2), gap_pressure (open concept-gap findings ÷ preference count, reused from doctor); records the vault_vitals metric
o2b brain benchmark           run --dataset <path> [--k N] [--expand] - hit@k + MRR against the live hybrid recall; records the recall_benchmark metric
o2b brain tune                run --dataset <path> [--k N] | status | reset - bounded self-tuning grid judged by the benchmark; persisted to Brain/search/tuning.json
o2b search <query> --expand   deterministic lex/vec/hyde expansion of a bare query (stopword-stripped lex, entity-context vec line, template hyde passage)

Wikilinks to frontmatter aliases: resolve at index materialization (schema v7): exact paths always win, a real basename is never shadowed, collisions resolve first-wins by sorted path, and o2b search status counts the pass via IndexStats.aliasResolved. Bridge discovery and clusters also run as maintenance-lane tasks after reindex. Self-tuning only changes behavior under search_self_tuning_enabled (or OPEN_SECOND_BRAIN_SEARCH_SELF_TUNING=1); an explicit --expand/expand always wins over the tuned default. Every surface appends one run-level record to Brain/metrics/<surface>.jsonl - the dashboard data contract documented in docs/metrics.md.

Write-path integrity and store safety (since v1.32.0)

o2b brain pending             list [--lane signals|notes|ingest|all] | apply <id> [--dry-run] | reject <id> --reason <text> [--dry-run] - review the write-approval queues (`write_approval.enabled` plus the per-lane keys); signals stage flat in Brain/pending/ (sig- ids), note creates under Brain/pending/notes/ (note- ids carrying the encoded publish target), ingest summary pages under Brain/pending/ingest/ (ing- ids); list shows every lane sorted and names unreadable entries with a reason, apply moves the unchanged document to its decoded publish target (--dry-run previews and writes nothing), reject renders it into Brain/retired/ with the reason; everything staged under Brain/pending/ stays out of the search index (admission reason `review-pending`), so recall cannot surface an unreviewed document
o2b brain signal retire       <id> --reason <text> [--superseded-by <id>] - move an inbox signal to Brain/retired/ with retire frontmatter (_status, retired_at, retired_reason, optional superseded_by, old-id alias); retired signals leave dream intake but stay queryable
o2b brain entity prune        [--confirm] [--json] - list entity nodes whose labels fail the structural quality gate (dry-run default); --confirm removes nodes and their edges behind the snapshot gate and reports the recovery point
o2b brain forget-source       --confirm now snapshots Brain/ before any deletion and reports the snapshot run id; dry runs take no snapshot
o2b brain doctor              gains the `symlink-escape` error (vault-internal symlink resolving outside the vault root) and the `entity-label-malformed` warning (prune candidates); `--remediate` gains the auto-safe `harden-permissions` step (owner-only chmod for existing Brain/ files, idempotent, step-capped, dry-run first)

Extracted facts pass a deterministic durability gate before persisting: structural signals only (temp paths, progress counters, run-id and timestamp shapes, measurement-token dominance, exit-status shapes), extendable with durability.denylist regexes; rejections log durability-skip events and are surfaced by count in the route result. Entity labels are sanitized (**Foo** becomes Foo) and structurally validated at every intake boundary, with entities.label_denylist as the only vocabulary source. Feedback writes that closely resemble a confirmed same-scope preference return a conflict advisory (and log write-conflict-advisory) while the write proceeds. Embedding-side: embedding_prefix_query/embedding_prefix_passage (env twins, e5 preset defaults) control instruction prefixes, invalid vectors are rejected with EMBEDDING_INVALID_VECTOR, and quota exhaustion surfaces as non-retriable EMBEDDING_QUOTA_EXHAUSTED with Retry-After-aware backoff for plain rate limits.

Belief lifecycle and decision memory (since v1.33.0)

o2b brain lifecycle           tombstone <path> --reason <r> | supersede <path> --by <path> | temporal-replace <pred> <succ> --at <instant> | tip <path> | curator [--slice <s>] - cross-type soft-delete and supersession: tombstone is idempotent frontmatter (file stays for audit, leaves recall/inject/active.md), temporal-replace closes and opens at one shared instant with half-open [valid_from, valid_to) intervals, tip resolves the supersedes chain, curator lists injected-never-used / contradicted / high-used memories
o2b brain claims              [--at <instant>] [--history] [--replaced <id>] [--contests <id>] [--rebuild] - claim-graph queries over existing relations and validity fields; current truth by default, history opt-in; --rebuild persists Brain/claim-graph.json deterministically
o2b brain decision            record --title <t> --chosen <c> [--assumption <a>] [--review-date <d>] [--premortem <p>] [--commitment <tier>] | outcome <slug> <text> | rate <slug> <1-5> [--rationale <r>] | show | list [--rated] | compare <slug...> | similar --title <t> | history [--subject <id>] [--cursor <c>] | recall --prompt <p> [--turn <n>] [--count <n>] [--last-turn <n>] [--surfaced-ids <ids>] - decision records under Brain/decisions/; record opens one review obligation per review_date, history pages decision_change.v1 receipts, recall is governed by decision_recall.max_per_session and decision_recall.min_spacing_turns; open decisions (write-side-trust): open --title <t> --question <q> --option <o> [--option <o>...] [--context <c>] parks a question at Brain/decisions/open-<slug>.md (duplicate question refuses naming the existing id), list_open [--status open|resolved|discarded] lists with unreadable records named, show_open <id> reads one, resolve <id> --choice <c> [--rationale <r>] mints the real decision page and stamps [[decision-<slug>]] plus one open-resolved receipt, discard <id> --reason <r> closes without deciding; terminal records stay in place
o2b brain tension             detect [--jaccard <n>] | list [--unresolved] | show <id> | confirm <id> | dismiss <id> | resolve <id> - persisted contradictions under Brain/tensions/ with an open -> confirmed/dismissed/resolved state machine; re-detection refreshes the existing note; unresolved tensions warn at context-pack build time
o2b brain tension             verify [<slug>] - read-only advisory decision-model verdict (contradicts | compatible | unrelated) per tension, or for every unresolved one; needs the optional `tension` use, else `available: false`; never changes a tension (see docs/decision-models/dedup-tension.md)
o2b brain authored-at-backfill  [--apply] - stamp authored_at on pre-1.33.0 session signals from their preserved turn instant; dry-run default, idempotent, never re-embeds
o2b brain session-grep        gains --since <time> and --before <time> bounds on turn time

Supersession is now consumer-aware: context packs inject only the tip of a supersedes chain under budget (an explicit historical flag keeps the whole chain), recall annotates superseded hits with their replacement, and the dream pass accelerates decay of low-recall superseded ancestors (chain-decay events). Decisions carry an optional commitment tier (exploring | leaning | decided | locked, also valid on preferences and theses) that renders in injected text in place of the raw confidence float when set. Every decision mutation and lifecycle transition appends an accountable decision_change.v1 receipt (before, after, evidence, confidence delta, actor, reason code) with durable idempotency keys. Search results expose authored_at when the source turn carried an instant, and exact hybrid-score ties order newer-first.

Source pipeline integrity and operator tooling (since v1.34.0)

o2b brain batch-plan          gains --src-subpath <rel> (scope a monorepo ingest to one subtree; escaping the source root is a typed error, and one crossing a submodule or nested checkout is refused outright; a subpath below a directory the repository itself ignores plans nothing and says so as an ignore warning), --exclude <pattern> (gitignore-style patterns through the shared ignore engine), and --reconcile (diff the plan's dispatched set against checkpoint completions and name every silently-lost source; a complete plan reports an explicit empty gap); a changed extraction contract is reported rather than absorbed - byte-identical sources reprocess once with the distinct contract-changed status (contract_changed and contract_changed_files keys under --json, an info line in text), a v1 manifest with no recorded contract included
o2b brain pre-extract         <file> [--json] - deterministic no-LLM code-structure extraction (classes, functions, imports, inheritance edges, and in .tsx/.jsx files uses edges to the components opening tags name, as JSON entity/edge seeds) for TypeScript, JavaScript, and Python, and Terraform (.tf, .tfvars) block seeds in address syntax (resource, data, module, variable, output, provider, locals) with depends_on and references edges and a module source as an imports seed; .tfvars yields variable names, never values; credentials in any import specifier are redacted (the userinfo of an http(s), git+http(s) or git:: specifier, a user:password pair or a bare token, and the value of a named credential query parameter such as sshkey, token or an S3 or GCS signing key, also on a specifier that is not a URL such as git@host:org/repo?sshkey=...; a conventional login such as ssh://git@ is kept, and a credential in a path segment, a fragment or an unnamed query parameter is not recognised); a specifier longer than 2048 characters is carried as the redaction placeholder; on the ingest path a source larger than 1 MiB is skipped with a reason naming the limit; relative imports are bound to files (resolved_to) only on the ingest path, which has the content manifest; unknown extensions report extracted:false with a reason, never an empty success. Terraform limits, by name: .hcl and .tf.json files, attribute values, nested blocks (lifecycle, dynamic, provisioner, connection) as entities, depends_on expressions, citations inside heredocs or template files, for_each and count expansion, moved, import, check and removed blocks, terraform {} settings and required_providers, resolution of a local module source to a directory, and a block header split across lines are not read, the items of a multi-line depends_on list are read as ordinary citations and seeded as references, not depends_on, and quotes inside an interpolation (`"${f("x")}"`) are not tracked, which can end it early and drop citations later on that line
o2b brain scan-citations      [--strict] - promote structural [Source: <name>, YYYY-MM-DD] markers in note prose into dated source-citation events on the temporal timeline; dedup on normalized name + date, malformed markers are reported and skipped (--strict exits nonzero on them)
o2b brain doctor              gains --repair (preview targeted fixes for doctor-detected classes: dangling workrun checkpoints, dead evidence links; dry-run default) and --repair --apply (perform the fixes idempotently, one typed doctor-repair event each); unfixable classes are reported needs-review; plain doctor and --strict stay read-only
o2b brain status              one consolidated operator snapshot: doctor, semantic health, hygiene, stale scan, review queue depth, active profile, and state-file health; every problem line carries the exact next command from the shared diagnostics-signal registry; healthy vaults print a compact all-clear

The hygiene file scan now honors nested .gitignore files with git semantics (a deeper ignore file scopes its own subtree, a nearer ! re-include wins, .git/info/exclude participates) through the shared path-scope engine also used by ingest scoping; repositories without ignore files scan exactly as before, and malformed patterns warn explicitly. Page discovery enforces the schema extractable allowlist: with a non-empty allowlist, pages whose schema_type is not listed are skipped up front and reported with a reason (an empty allowlist changes nothing). Both o2b and vault-log treat an early-closed stdout pipe (o2b ... | head) as a clean exit 0, while any other stdout error now fails loudly.

Ingest discovery applies the repository's own declarations to --src-subpath as well: the planner descends from the source directory to the subpath root and layers each intermediate directory's .gitignore on the way, so a scoped plan answers the same about a tree as an unscoped one. A subpath that starts below a directory the repository ignores is not walked: the plan is empty and carries an ignore warning whose source is --src-subpath, naming the ignored directory and pointing at the --exclude ! re-include that opens it. A subpath that crosses a submodule or nested checkout is refused with an error naming the boundary directory instead: those files belong to that repository, no --exclude pattern can re-open a boundary, and the error says to plan against that repository directly. Ignore warnings print beneath the plan on the human surface and travel as the ignore_warnings array (source, line, pattern, reason per entry) under --json; line is 0 for a warning about no single pattern line, which covers both a --src-subpath warning and an ignore file that could not be read.

Parallel ingest workers share the content manifest, the plan checkpoint, the session ledger and the git record store, and each write to one of them waits for a file lock. The wait is 5 000 ms by default; set OPEN_SECOND_BRAIN_LOCK_WAIT_MS (a whole number of milliseconds, 0 means one attempt) to give a slow host (antivirus scanning, network or synced folders) a longer one. A value that is not a whole number is an error, not the default. The variable does not change the one-second wait of interactive commands such as the architect run. Waiters take turns: a writer that releases a lock someone is waiting on hands it over before it takes it again. A write that still cannot get the lock is refused with ELOCKED and lock busy: <lock file>, writes nothing, and says what to do next: retry the ingest, run the parallel ingests one at a time, or raise the wait.

Trusted recall and memory write surface (since v1.35.0)

o2b doctor                    gains --readiness: six functional probes (model-inference key resolvable, embedding provider loadable with model and dims, runtime-adapter construction, installed runtimes verified off disk through each adapter's own verify; since v1.64.0 every registered client command still resolvable - a proved-absent absolute path fails naming the client and `o2b install <target> --apply`, a relative path is unresolved because the client resolves it from its own working directory, an unresolvable bare name is unknown because the host spawns with its own PATH - and the workspace agent-instruction files (`AGENTS.md`, `CLAUDE.md`, `GEMINI.md`) carrying the marker write-back contract block, conforming / missing-block / malformed-block / missing-clauses / absent / symlink / unreadable; a file without the block is graded skipped because the block is not installed, while a malformed or incomplete block fails) with per-check timeouts and outcomes pass, fail with a reason, skipped-not-configured, or unknown-could-not-measure; a failure and an unmeasured probe exit with different non-zero codes (see "`o2b doctor` exit codes" below); without the flag output stays byte-identical
o2b brain morning-brief       renders recalled items as one chronological Recent activity timeline with a per-item structural type marker and a relative age label; the underlying JSON data arrays are unchanged

Prompt-time recall is a hook, not a verb: with recall_inject_enabled: "true" (env OPEN_SECOND_BRAIN_RECALL_INJECT_ENABLED) the UserPromptSubmit hook injects a bounded brief of relevant vault notes (by default 4 notes, 900 chars, a 2,500 ms time budget and a 0.35 confidence floor, each tunable as described below), fenced as untrusted content with neutralized titles; any internal error or timeout injects nothing, and every decision writes one audit line recording counts and scores only. The knowledge-gap loop is likewise hook-driven: with gap_loop_enabled: "true" (env OPEN_SECOND_BRAIN_GAP_LOOP_ENABLED, threshold via gap_loop_threshold), recurring recall gaps promote at session end into durable task notes under Brain/gap-tasks/ (stable-key dedup, never the kanban board), render as a session-start agenda, and auto-close once the topic is later recalled confidently.

Recall-inject tuning keys live in the flat global config next to recall_inject_enabled and matter only while that flag is on. Each one has an env override, and the env value always wins over the config value:

Config key Env override Range Default
recall_inject_max_notes OPEN_SECOND_BRAIN_RECALL_INJECT_MAX_NOTES integer 1..10 4
recall_inject_max_chars OPEN_SECOND_BRAIN_RECALL_INJECT_MAX_CHARS integer 200..8000 900
recall_inject_time_budget_ms OPEN_SECOND_BRAIN_RECALL_INJECT_TIME_BUDGET_MS integer 250..6000 2500
recall_inject_confidence_floor OPEN_SECOND_BRAIN_RECALL_INJECT_CONFIDENCE_FLOOR number 0..1 0.35
recall_inject_dedupe OPEN_SECOND_BRAIN_RECALL_INJECT_DEDUPE boolean true

The four caps resolve leniently: a value that is out of range, non-numeric or (for the integer caps) fractional keeps the built-in default, and the hook names the rejected key (or its env variable when the value came from the env) in its audit line (config_invalid). Unset caps leave the brief byte-identical. recall_inject_dedupe is on unless set to the literal "false" or "0". With it on and a host that sends a session_id, a note span this session was already shown, by an earlier brief or by the SessionStart digest, is not injected again; a host without a session id gets no dedupe. The semantics are in the recall-inject decision-model page.

Recall slices are vault policy, not machine config: a recall_inject: block in Brain/_brain.yaml declares named slices, each retrieved with its own filters and rendered under its own heading inside the one fenced brief. The block is deliberately absent from the generated _brain.yaml template, because slices name this vault's folders and note types and have no default. Without the block the hook keeps its single unsliced relevance query, byte for byte. Contract example:

recall_inject:
  slices: [decisions, lessons]
  slice_decisions_heading: Recent decisions
  slice_decisions_path_prefix: Brain/decisions/
  slice_decisions_types: [decision]
  slice_decisions_limit: 2
  slice_decisions_max_chars: 400
  slice_lessons_path_prefix: Brain/lessons

slices lists the slice names in the order they are laid out; a name matches ^[a-z][a-z0-9]{0,23}$ (no underscore, so each slice_<name>_<field> key splits unambiguously), and at most 6 slices are allowed. The per-slice fields are heading (defaults to the name), path_prefix (vault-relative), types (an inline array matched against frontmatter type; empty means no class filter), limit (integer 1..10) and max_chars (integer 100..8000). A slice with neither limit nor max_chars takes the global caps. An unknown field for a declared slice, a slice_* key for an undeclared name, a duplicate name, more than 6 slices, an out-of-range number or a path_prefix with .., a leading / or a drive letter is a hard load error. Every slice is retrieved in parallel under the one shared time budget, so more slices trade against latency on a large vault. A _brain.yaml that fails to load takes no slice path and is recorded as slices_config: "invalid" on the audit line; because the search reads the same policy file, the decision then ends in error with fault retriever_failed and nothing is injected until the file is fixed.

Search-side trust switches: search_trust_gate_enabled (env OPEN_SECOND_BRAIN_SEARCH_TRUST_GATE) zero-ranks quarantined material (self-approval quarantine, untrusted-source provenance, entity contamination) out of both semantic and lexical results and attaches the memory_trust_assessment and retrieval_decision_trace receipts naming every exclusion; search_supersede_fade_enabled (env OPEN_SECOND_BRAIN_SEARCH_SUPERSEDE_FADE) makes an inbound supersedes / superseded_by relation on a successor fade the unchanged superseded note by a named multiplier. Both default off and leave ranking byte-identical when unset.

Knowledge intake and consolidation (since v1.36.0)

o2b brain telegram-capture   run | catchup - explicit inbound capture runner (fetch-based long-poll getUpdates); gated by telegram_bot_token (env TELEGRAM_BOT_TOKEN) and the telegram_chat_allowlist chat-id allowlist (env TELEGRAM_CHAT_ALLOWLIST, empty accepts nothing); each accepted message becomes one staged capture note under Brain/captures/ with provenance; rejected updates log one decision each; catchup replays captures since the last acknowledged one; a missing token is a typed startup error
o2b brain inbox-drain        [--apply] - walk staged captures, classify each structurally (URL-shaped body -> source reference, explicit obligation marker -> task, otherwise atomic idea), route on apply (source ingest, note in captured/, obligation open), archive processed captures, and report every item with action and reason; dry-run default writes nothing; rerun after apply is a no-op. Since v1.64.0 an applied idea route also stages its page's area-hub outcome in the repair lane's apply-gated store - an `area_membership` candidate for exactly one structurally validated hub, a `skip-no-hub` refusal for none, `skip-ambiguous-hub` listing every candidate for several - and a store failure aborts the drain loudly before the archive, so the still-staged capture makes a rerun recover it
o2b brain diarize            <entity> [--json] - entity profile skeleton plus one needs-llm-step envelope; the stated-vs-evidenced section is computed deterministically (stated claims versus evidence frequency and recency), every line carrying an evidence identity; unknown entity is a typed error
o2b brain repair-lane        [--apply --confirm "apply repair"] - propose link-graph edges ordered by identity strength (explicit references, session continuity, same-topic evidence; inferred opt-in) under a confidence threshold and a hard per-run write cap; dry-run default; reruns after apply converge to zero writes. --apply plans first and then runs the paired graph-efficacy holdout harness over every edge that plan proposes, taking anchor and target from the edge itself: it reports graph lift (targets reachable only through the graph) apart from direct recall (targets already a 1-hop neighbour) and refuses the apply, writing nothing, when any target resolves to no durable-memory note or hydrates into no evidence. The refusal names the failing edges with their verdict and resolves its exit through the registered repair-holdout-unresolved diagnostic; --json carries the counts in an additive holdout object beside next_command. Dry-run does not evaluate the gate and its bytes are unchanged - it already names each unresolvable endpoint as a skip-missing-target decision. Explicit references pool page titles and frontmatter aliases; a mentioned term that exactly one page carries proposes an edge, and a term several pages carry proposes none - each (mentioning page, carrying page) pair is reported as a skip-ambiguous decision, never written, never counted against the write cap and never sent to the holdout gate. Since v1.64.0 the lane also merges the hub candidates and refusals `o2b brain inbox-drain` staged at intake (Brain/.state/repair-candidates.jsonl) with the graph-collected ones before planning, so a hub proposal made when a capture was routed reaches the same apply + confirm + holdout gate; refusals propose no edge and are never holdouts. The store keeps one record per routed page, prunes records whose page is gone, is reported by `o2b state status` as the `repair_candidates` surface, and a corrupt line is dropped and counted as `skipped_corrupt` on the drain report
o2b brain orphan-repair      [--apply --confirm "apply orphan repair"] - detach dangling session references from observation signals (the findings `o2b brain doctor` reports as `orphan-session-ref`, each naming this verb on its `fix` field); --apply removes ONLY the `session_ref` key from each signal's frontmatter, keeping the observation body, topic and every other field, and quotes the detached value in the report; dry-run default writes nothing, the exact confirmation phrase is required for apply, a hard per-run write cap bounds the run, and a rerun after apply converges to zero writes; the doctor pass never repairs
o2b brain design-note        <topic> [--payload <json> | --payload-file <path>] [--agent <name>] [--json] - the one-shot sibling of `o2b brain panel`. Without a payload it is read-only: it grounds the topic in the vault's tension records, decision records and truth projections and prints the single needs-llm-step envelope the calling agent answers, naming any store the vault holds nothing in (which is not the same as a store that matched nothing). With a payload it validates the written note and commits it as Brain/decisions/design-<date>-<topic>.md. The note must weigh named alternatives and mark EXACTLY ONE recommended: zero and two-plus are both refused, and the refusal states the count. A second note for the same topic on the same day is refused, never overwritten
o2b brain skill-proposals    page-candidates [--json] - read-only: gate the vault's user pages on the page-meta trio (core tier, non-stale lifecycle, high confidence) and an observed-reuse floor, skip any page an installed skill already covers, and return one needs-llm-step envelope per admitted page plus every skip with its reason
o2b brain skill-proposals    page-draft <page> (--payload <json> | --payload-file <path>) [--json] - validate a returned SKILL.md draft and stage it as a pending mature_page proposal INSIDE the vault; accept is what materializes the SKILL.md under the configured skills root, through the write-ahead journal
o2b brain extract-signals    <session-ref> [--payload <json> | --payload-file <path>] [--agent <name>] [--json] - mine durable taste signals from an already-imported session's user turns. Without a payload it is read-only: it prints the turns it would mine and the single needs-llm-step envelope the calling agent answers. With a payload it validates the answer and writes the accepted items into Brain/inbox/ as speculative source_type: auto_extract signals, subject to the durability denylist and to Brain/pending/ staging when write approval is on. A payload over the per-session cap, an item below the confidence floor, or two items sharing a topic refuses the whole payload by name and writes nothing; a session with no imported turns is refused, never reported as empty
                             With the optional decision-model turn pre-filter (use extract_prefilter, docs/decision-models/extract-prefilter.md) the --json plan may add turns_dropped, skipped { reason: decision_model_prefilter, turns_dropped } with llm_step: null (nothing to mine), and decision_model { degraded }; none appears with the use off. Payload items accept an optional string source_turn
                             Since v1.74.0 every turns_mined entry carries timestamp, the turn's stored timestamp verbatim ("" when the source has none), and each prompt line reads [<turn-id> @ <timestamp>] <text>, or [<turn-id>] <text> for an undated turn - the run's clock is never substituted. The envelope instruction holds the items to five language-neutral rules: a time bound is written as an ISO 8601 date or interval resolved against the stating turn's timestamp (an end-only bound as an interval, because the dream pass reads a lone date as a start), conversational mechanics are skipped, one rule per item, a rule's conditions stay in its principle, and restatements are dropped. The dream pass turns the ISO bound into the preference's valid_from / valid_until through its existing temporal extraction; the extract lane writes neither field itself. Two items with the same topic refuse the payload under semantic_cross_item, the message naming both indices (items[i] and items[j] share topic "<topic>"), before any item is hashed or written
o2b brain deep-synthesis     --json now also carries findings with causal_context, decomposed confidence (support, opposition, freshness, coverage), and the excluded_findings ledger with excluded_finding_count

Web research gains keyed providers: Brave and Tavily join the research pool only when BRAVE_API_KEY or TAVILY_API_KEY is set, through a shared external-fetch helper (typed network, auth, http, and payload errors; response cache keyed by the normalized request; keys never appear in cache keys or error text). A keyless pool reports itself empty and report writing stays byte-identical. The full-page extract step feeds verbatim page text into the existing citation-constrained pipeline. The dream pass gains a count-triggered fact rollup ladder: an optional rollup: config block (fact_threshold, identity_threshold) makes a tier that accumulates enough new facts emit one needs-llm-step rollup envelope over a persisted ledger (Brain/rollup-ladder.json); below the threshold dream output is byte-identical. Skill proposals pass a deterministic pre-promotion verifier (rejections recorded with reasons), carry a version that increments on evolution, and merge same-name collisions instead of forking.

Retrieval quality and context delivery (since v1.37.0)

o2b brain state              set <aspect> <value> | get <aspect> | list | clear <aspect> - overwrite-only per-aspect operational-state lane under Brain/state/; writing an aspect replaces its canonical value; the lane never enters FTS/vector/graph and a retrieval-time barrier drops superseded exact-state rows from recall; Brain/pinned.md deliberately stays a searchable scratchpad
o2b search rerank-fit        sample recorded demand-log queries and correlate reranker scores against the base retrieval signal; verdicts fits | out_of_domain | inverted | inapplicable with a disable-or-swap recommendation; strictly read-only (probes disable cache and self-heal)

Relational retrieval: search_relational_arm_enabled (env OPEN_SECOND_BRAIN_SEARCH_RELATIONAL_ARM, default off) adds a fourth RRF arm that parses relationship-shaped queries against the schema-pack link_types vocabulary (subset validation, no word lists) and runs a bounded depth-2 typed-edge fan-out; RRF and dedup keys carry source identity and the query cache scopes by canonical source-set key. Off means byte-identical ranking. Query plans also carry a surface verdict that routes structurally summary-shaped questions (source-targeted, artifact-kind, summary-typed pages) to the summary surface; non-summary routing is unchanged. Search accepts optional session/project scope filters (MCP brain_search session_scope / project_scope), and page dedup keys fold composite scopes additively so identical text in two scopes stays distinct.

Context delivery: nav_tier_enabled (env OPEN_SECOND_BRAIN_NAV_TIER_ENABLED, default off) injects an additive, untrusted-fenced navigation tier (top hubs by link degree plus vault size) on UserPromptSubmit when the nav_tier_cadence_minutes cadence (env OPEN_SECOND_BRAIN_NAV_TIER_CADENCE_MINUTES, default 30) is due; every inclusion decision is audited with reason and added characters. hook_strict_enabled (env OPEN_SECOND_BRAIN_HOOK_STRICT_ENABLED, default off) makes the first raw vault-file read of a session receive a one-time deny that names the brain search surface, then downgrades to a soft nudge; any brain query/search refreshes an orientation stamp that suppresses the block, and every failure path fails open. Hook stamps live under .open-second-brain/hook-state/ with epoch-ms expiry. hygiene_digest_enabled (env OPEN_SECOND_BRAIN_HYGIENE_DIGEST_ENABLED, default off) adds an end-of-turn Stop-hook line that surfaces pending Brain hygiene findings once per change, folding warning- and action-severity findings into per-detector counts and staying silent while the findings set is unchanged. Claude Code only: runtimes whose Stop block would force a continuation turn stay silent. o2b partner codegraph report and o2b doctor aggregate codegraph status across every discovered code project, threading project_path per query when supported (feature-detected) and degrading with an explicit note when not.

Chunked re-grounding: reground_parts_enabled (env OPEN_SECOND_BRAIN_REGROUND_PARTS_ENABLED, default off) splits an oversized SessionStart payload instead of emitting it whole. Claude Code persists an additionalContext past roughly 10,000 UTF-16 units to a file and shows only a preview, so with the flag on, a joined standing-rules, scoped-rules and memory payload longer than the part ceiling is cut into at most 8 parts (at block, then paragraph, then line boundaries, in priority order). Each part opens with [Open Second Brain context - part i of n] and every part but the last closes with (continued in part i+1 of n); when content past the 8th part is dropped, the last part closes with (context truncated: N further part(s) not delivered) instead. Part 1 is emitted at SessionStart and parts 2..n are queued in the session's hook state; the reground-deliver hook then hands out exactly one queued part per PostToolUse or UserPromptSubmit event until the queue is empty or the next SessionStart replaces it. A tool call inside a delegated sub-agent (the payload carries agent_id) takes no part, so the queue stays for the main agent. Only a SessionStart event splits and starts a new queue; a run of the SessionStart hook on another event emits the whole payload and leaves the queue alone. Only Claude Code and Codex payloads with a session id are split, because only they have the carrier registered; every other runtime, and any payload that fits the ceiling, gets the single payload as before. The ceiling is in UTF-16 code units:

Config key Env override Range Default
reground_part_chars OPEN_SECOND_BRAIN_REGROUND_PART_CHARS integer 2000..100000 9000
reground_part_chars_claudecode OPEN_SECOND_BRAIN_REGROUND_PART_CHARS_CLAUDECODE integer 2000..100000 reground_part_chars
reground_part_chars_codex OPEN_SECOND_BRAIN_REGROUND_PART_CHARS_CODEX integer 2000..100000 reground_part_chars

The runtime key wins over reground_part_chars, which wins over the default 9000 (the observed Claude Code threshold less 10%, applied to Codex too until measured). At each level the env value wins over the config value, and an invalid value falls through to the next level and is named in the receipt's config_invalid (by its env variable when the rejected value came from the env). The queue lives under .open-second-brain/hook-state/ with a 24 h expiry, in a per-session file written with mode 0600; when .open-second-brain or hook-state is a symbolic link, the hooks neither read nor write through it. The receipt and audit fields it adds are listed in observability.

Semantic-health baselining (since v1.38.0)

o2b brain health-baseline    set <date>|now | get | clear - record, show, or remove the health.silence_before watermark in _brain.yaml without hand-editing it; set accepts a date-only or full ISO-8601 value (or the literal now) and rejects anything else with exit 2; the upsert preserves unrelated config content byte-for-byte under a file lock with an atomic rename

The watermark acknowledges vault history up to a date: o2b brain health suppresses advisory findings whose underlying entries are entirely older than health.silence_before - a batch-concept-inflation burst whose window ended before it, a concept-gap term whose corpus mentions all predate it - and computes the verdict from the surfaced findings only. An entry without a parseable timestamp counts as newer than any watermark, so it is never suppressed, and a concept-gap term with even one post-baseline mention surfaces with its full frequency. Whenever the watermark hides at least one finding, the report prints suppressed: N finding(s) older than baseline <date> and carries an additive suppressed object with per-detector counts; nothing is ever hidden silently. An invalid config value is an explicit error at load, and with no watermark set every output is byte-identical to v1.37.0.

Context integrity gates (since v1.39.0)

o2b search query --agent-scope <name>    restrict results to pages that are ownerless or owned by <name>
o2b search expand --agent-scope <name>   same rule on the chunk drill-down; a withheld chunk is indistinguishable from an absent one
o2b brain sgrep --agent-scope <name>     same rule on the structured grep surface
bun run link-ratchet                     measure the vault-wide broken-link count and record it as the new ceiling
bun run link-ratchet:check               fail with exit 1 when the count has risen above the committed ceiling

Three gates live in an optional integrity: block in _brain.yaml: owner_scope_delivery (default off), embedding_abi (default warn), and pack_validity_seconds (default 900). The two gate keys take off, warn, or fail; an unrecognised mode or a non-positive validity is an explicit config error at load, never a clamped default.

owner_scope_delivery governs the preference-backed delivery surfaces - the context pack, the pre-compress pack, the morning brief, the digest, and active.md. The search-backed surfaces above filter whenever a scope is passed and are not gated. Ownership fails closed: a page whose frontmatter cannot be read, and a page whose owner: is present but not a plain string, are withheld from a scoped caller rather than treated as shared. Retiring a memory keeps its owner.

embedding_abi compares the recorded embedding model, dimension, and sqlite-vec version against the running ones on every read-mode open. o2b search check reports a drift with a copy-pasteable fix regardless of the gate, because it is a diagnostic the operator ran deliberately and refuses nothing; o2b search status reports it under warn; under fail the read-mode open raises a named error. A store predating the stamp is reported, never refused, and nothing is rebuilt automatically - the self-heal path rebuilds keyword data only and would leave stale vectors in place.

The broken-link ratchet counts link rows the read-time resolution ladder cannot resolve and compares them against link-ratchet.json at the repository root. --check is the same code path with writing disabled, so detection is identical between the two forms; it exits 1 on a rise and names the write form as the fix. A subject that measures nothing - empty, missing, or fully excluded - is an explicit unmeasurable state, not a ceiling of zero, and the count is refused unless the index's last run was a forced full pass.

o2b brain doctor gains an uncertain stream for conditions it attempted but cannot claim completed cleanly: frontmatter lines the parser dropped, lineage-ledger findings, stale lock files, and a missing vault identity marker. These render as [UNSURE] lines and appear in --json under uncertain; they do not affect the exit code. A vault with no Brain/ layer is now reported as a named brain-root-absent warning instead of clean, which means o2b brain doctor --strict exits 2 on an un-initialized vault where it previously exited 0.

Write binding (since v1.43.0)

An optional write_binding: block in Brain/_brain.yaml declares where a CALLER-NAMED write may land:

write_binding:
  path_prefixes:
    - Projects
    - Journal/Weekly

It covers the four tools that take a vault-relative path from the caller - brain_create_note, brain_update_note, brain_append_note and brain_write_batch - because all four resolve that path through one shared envelope. Slug-derived writes (a signal, a preference, a dead end) and fully derived ones (daily logs, dream runs, continuity records, every telemetry surface) compute their own destination and are out of scope: a prefix rule could only refuse those wholesale, which is an off switch rather than a boundary. tests/core/architecture/write-site-census.test.ts is the standing record of every in-vault write site the binding does not cover.

The check reads no identity at all. Agent names here are self-asserted and additionally overridable per call, so a fence keyed to one is bypassed by passing a different string; the binding's whole authority is _brain.yaml, which the operator controls and which no MCP call can rewrite. It is a write boundary over caller-named paths, never a security boundary and never per-credential.

Both the name and the realpath must be admitted, so a symlinked directory under a declared prefix cannot widen it. A refusal names the path the caller gave, where that name actually resolves when the two differ (or a location outside the vault), the prefixes in force, and the registered exit for write-binding-refused. An absent block, or a block without path_prefixes, is inert: every write path behaves byte-identically to before the key existed.

Materialization freshness (since v1.46.0)

o2b brain clusters run --if-stale recomputes only when the derived digests under Brain/clusters/ are not fresh. It used to answer that question with a boolean derived purely from file times: outputs newer than every input meant fresh, and anything else fell through to a recompute. Two states were folded into one by that shape. An output the walker could not stat looked exactly like an output that was simply out of date, and a materialization stayed "fresh" forever as long as nobody touched an input, however old the digest or however much the code that produced it had changed underneath it.

The verdict is now three-state, and each state is a distinct answer:

freshness Meaning freshness_reason
fresh outputs are current; --if-stale skips the run null
stale outputs must be recomputed not_materialized, input_newer, or ceiling_exceeded
unknown the measurement itself failed; the run recomputes AND names which half could not be read outputs_unreadable or inputs_unreadable

The stale reasons are: not_materialized (nothing has been written yet), input_newer (an input note is newer than the oldest output), and ceiling_exceeded (the new wall-clock rung, below). The whole freshness_reason vocabulary is closed - those five values plus null.

New config key health.materialize_max_age_days. A wall-clock ceiling on a derived artifact's age, default 30. Past it, --if-stale recomputes even when no input note has moved, so a materialization cannot outlive a release cycle by sitting untouched. It must be a positive integer; 0, a negative, a fraction, and a non-number are each refused at config load with health.materialize_max_age_days must be a positive integer rather than clamped to a default. An age exactly equal to the ceiling is not past it. Absent health: block, or absent key, means the default. A Brain/_brain.yaml that does not parse now fails the run loudly instead of quietly falling back and reporting a skip.

--json carries communities and skipped: "fresh" on a skip, and adds a staleness object (state, reason) on an unknown verdict, where the run proceeds but says what it could not measure. A plain stale verdict adds no key: the recompute is the answer. Every --if-stale run records the verdict to the communities metric surface (see metrics.md); before this only skips were recorded, so the runs that mattered were the ones nothing measured.

Recall channel coverage (since v1.46.0)

Recall telemetry records now carry a channel naming the seam the recall came through, from the closed set mcp, cli, hook. It is a fact about the emitting seam rather than an argument, so no caller can set it: the MCP handlers stamp mcp, o2b brain context-pack --telemetry stamps cli, and the prompt-time recall-inject hook stamps hook. Every emit site is required to name one, and that requirement is enforced when the code is built rather than at run time - there is no runtime rejection to observe, and no record is written without a channel.

Read it back with --channel on o2b brain recall-telemetry list and summary, and in the by_channel rollup that summary prints and carries under --json. by_channel omits a channel with no record rather than reporting it as 0, so absence of a bucket means "nothing measured here", never "measured zero".

Records written before v1.46.0 carry no channel. They still count toward total, by_mode, by_status, and the gap counts, but they land in no by_channel bucket and --channel <any> excludes them. Nothing is back-filled: a continuity record is historical and guessing its seam would be inventing the fact the field exists to record.

o2b brain doctor crosses those counts against whether a channel is installed at all, over a 30-day window, and reports two codes:

  • recall-channel-silent (warning) - the channel is installed and expected to deliver, and recorded nothing in the window. Installed and quiet is a different condition from never installed, and only the first is worth telling you about. Today only hook can raise it, because MCP and CLI telemetry are per-call opt-ins with no persisted install state: an absence there means nobody requested a record. Its next command is o2b brain recall-telemetry summary --channel <channel> --since <iso>.
  • recall-channel-unmeasured - one half of that cross could not be read: the recall-inject config gate would not resolve, the hook audit root denied the walk, or the telemetry records themselves could not be opened. It rides the doctor's uncertain stream rather than the issue streams, so it renders as an [UNSURE] line and does not affect the exit code, and it carries no next command by design - the finding is that a reading failed, and no single o2b invocation repairs that.

An installed channel that delivered at least once reports nothing, and a channel nothing asked to run reports nothing. The check is fail-soft: it can never fail the doctor pass it rides on.

Session logs this machine already has (since v1.50.0)

o2b brain import-session <path> needs a path, which made cross-agent recall worth nothing for any session nobody knew to import: an operator had to already know that Claude Code writes under ~/.claude/projects, that Codex has moved its rollouts between four subdirectories of $CODEX_HOME, that Grok encodes the working directory into a path component - and then run the verb once per file. Three flags are the machine-wide half, sweeping the roots src/core/runtime/host-facts.ts declares:

o2b brain import-session --status              Per-runtime found / imported / gap counts, and the roots each was looked for under
o2b brain import-session --discover            The same report, plus the absolute path of every log that has never been imported
o2b brain import-session --discover --all      Import that gap, one file at a time, through the same importSession the path form calls
o2b brain import-session --status|--discover --progress
                                               Newline-delimited progress on stderr under the `sessions` operation, one `hash` stage

The modes never blend, and a conflicting combination is refused by name rather than ranked: a path with --discover or --status, --discover with --status, or --all without --discover. Either ranking would be a guess about which one the operator meant.

--discover --all writes through the same importSession the path form calls rather than a second copy of it, so redaction, the tool-payload exclusion and the dedup index are the ones that were already there. One unreadable log does not lose the rest of the sweep: the failure is named on stderr, the run carries on, and the exit is 1 if anything failed.

The ledger is keyed by content, and lives with the other derived state at <vault>/.open-second-brain/session-import-ledger.json - the session_import_ledger row of o2b state status. Each entry records the SHA-256 of the file's bytes at the moment it was imported, so a log whose bytes have not changed is reported as imported without anything opening it again; re-importing an unchanged log used to re-parse every byte to discover it had nothing to write, because dedup suppresses the WRITES and not the read. Content and not mtime: an rsync, a Syncthing round trip or a git checkout moves the clock and not a byte, and a ledger that believed the clock would re-import the whole machine after any of them. A named-path import records its file too, when that file sits under a declared root, so a later --status does not offer back what the operator already imported by hand. A corrupt or unknown-schema ledger is a hard error rather than a silent reset, which would report an operator's entire imported machine as outstanding.

A Cursor database is counted as unparsable, not as a gap. Cursor keeps per-workspace chat state in state.vscdb, a SQLite file this build can LOCATE and no adapter it ships can read. A gap is work an operator can close; an unreadable format is not, and folding the two would put a permanent number in a column an operator is trying to drive to zero. So the identity the report prints is found = imported + gap + unparsable, and the third term is visible in both the human lines and --json. Roots that exist and could not be listed, and files that could not be hashed, are carried separately again as unreadable, because a partial sweep found less than it should have and the counts alone cannot say so.

The sweep reads only to hash and writes nothing - not even the ledger, which records imports and would be lying if a report updated it. It runs under the sessions cooperative deadline (safeguard_timeout_sessions_seconds); see "Forward pointers" below for why that operation has a budget of its own.

The transcript corpus (since v1.50.0)

o2b brain export --format transcripts-jsonl --transcripts <file|dir>
                 [--runtime <adapter-id>] [--since <iso>] [--until <iso>]
                 [--out <path>] [--force]

One JSON object per line, one object per CONVERSATION: schema, runtime (the adapter the file was detected as), session_id (the transcript's basename, never its absolute path), started_at, ended_at, message_count, and the ordered messages.

The alternative - one instruction/response PAIR per line, the shape a supervised fine-tuning harness eats directly - was considered and rejected, and the reason is worth stating because picking either silently is how a guess becomes an authoritative fact downstream. A pair set is a lossy projection: producing it means deciding which turn is the instruction, what to do with system and tool turns, and where a multi-turn exchange breaks into pairs - three judgements the runtime never recorded and this exporter cannot recover. A harness that wants pairs can derive them from a conversation; nothing derives the conversation back from the pairs.

A message carries turn_id, role, timestamp, text, and the NAMES of the tools that turn called. Tool names travel and tool inputs never do: inputs are host paths, command lines and pasted payloads, and this is a conversation dataset rather than a tool-trace one. A turn with neither text nor a tool call - a runtime queue event, a payload-less meta line - is not a message and is counted rather than emitted.

--transcripts names a file or a directory; a directory is walked for *.jsonl with symlinks skipped rather than followed, and a tree nesting deeper than 32 directories is refused by name rather than overflowing the stack. A path naming a FILE is taken whatever its extension, because the operator named that file, and a level of the walk that cannot be listed refuses too - an unreadable source is not an empty one. --runtime keeps one adapter's transcripts and is resolved for its refusal, so an unknown id lists the registered set instead of reporting an empty corpus.

--since / --until select whole conversations by their START, so a kept conversation is never sliced at the window edge and the rest of a rejected file is never read. --until is exclusive. A first turn whose timestamp cannot be read refuses the export rather than guessing which side of the window it falls on - including the case that actually happens, where the runtime recorded no clock and the adapter reports the epoch, which sorts before any window an operator would type.

The corpus never exists whole in memory. Records are guarded and written one at a time to a spool file the operator never named, then rename(2)-d into --out or streamed to stdout a chunk at a time; peak memory is one conversation plus one chunk.

Refusals. A *.jsonl under the named source that no adapter recognises stops the export, naming the file: a corpus that quietly omits a transcript reads exactly like a machine that never recorded it. A ZERO-BYTE file is the one exception and is counted (empty) instead - it carries nothing to omit, Claude Code leaves them behind whenever a session is killed before its first turn flushes, and refusing on them made the export unusable against a real store. And a secret-shaped identifier refuses the whole export and writes nothing: the spool is deleted and no file the operator can see was ever touched. The refusal names the conversation by its basename - except when the basename IS the secret-shaped identifier, where it is named by runtime and start instant instead, because printing it is the one thing that code path exists to prevent.

An export that matched nothing says which filter emptied it, on stderr, with every scanned file on exactly one counter:

note: no conversation matched; <scanned> transcript file(s) scanned, <n> from another runtime, <n> outside the window, <n> with no exportable turn, <n> empty

The four reasons sum to scanned minus the exported records, so an empty file under exit 0 can never be read as a machine that recorded nothing.

Source distillation

o2b brain distill            <source> (--claims <json> | --claims-file <path>) [--strict-quotes] [--excerpt-file <path>] [--agent <name>] [--vault <path>] [--json]

Writes one idempotent distillation page per source from atomic claims the calling agent supplies as a JSON array of { "text": "...", "block": "^abc" } objects (or an object with a claims array), at most 1000 per call, each claim's text on one line. <source> is a vault-relative path or a URL. Open Second Brain runs no model: it validates the claims, checks every quoted span in them against the source, and writes the page.

Since v1.67.0, every quoted span in a claim is compared with the block the claim cites, or with the whole source when the claim cites no block. A quotation mark is any character with the Unicode Quotation_Mark property, paired by position, and an apostrophe inside a word is never a delimiter. The comparison normalises both sides the same way (NFC, inline Markdown reduced to display text, block markers and a trailing ^id removed, whitespace collapsed) and never folds case, punctuation or wording; an ellipsis splits a span into fragments that must appear in order. A cited block that does not resolve fails the span rather than falling back to the whole source. The full rules are in the MCP reference.

  • A span that fails is unquoted: its quotation marks are removed, the words stay, the page is written, and the result reports it.
  • --strict-quotes refuses the whole write instead, exits 1, and writes nothing. The message reads distill: quoted spans failed verification: followed by claim <index>: <outcome> pairs, never claim text.
  • Since v1.67.0, --excerpt-file <path> stores the verbatim text read from a source the vault does not hold (a URL, or a path with no file). The page records capture_scope: bounded-local and an excerpt_hash, keeps the text under a ## Excerpt heading, and checks quotes against it. The excerpt is refused, before anything is written, for a source the vault holds, when it holds no text or contains NUL, or above 65,536 bytes. A file whose bytes are not valid UTF-8 is a usage error (exit 2, distill: excerpt file is not valid UTF-8), so the stored excerpt is always the file's own bytes.

The success line keeps its earlier form for a clean run over a local source and gains suffixes in this order: [untrusted_source] when the page is in the untrusted lane, [bounded-local] when an excerpt was stored, and [quotes verified:V unquoted:U] when spans were checked, where V is the number of spans verified in a block or in the source and U the number unquoted. --json adds capture_scope (always: full-local, bounded-local or url-only) and, when spans were checked, quotes: checked, verified_in_block, verified_in_source, unquoted, unpaired, findings (each { claim, outcome, span }, capped at 25 with total, returned and truncated).

Non-Markdown sources (since v1.69.0)

o2b brain extract             <file> [--json] - read-only preview of what ingest derives from a file, chosen by its extension; writes nothing
o2b brain batch-plan          plans .csv .tsv .html .htm by default; a planned file of a non-text format shows it: "- <path> (<status>, <bytes>B, <format>)"; PDF, Office, EPUB, RTF and image files are listed as format-not-extractable with their format instead of counted as unclassifiable

o2b brain extract reads one file (a path on disk, inside a vault or not) up to 8 MiB and prints what brain_ingest_source would derive from it. For HTML it prints the title and one line per part:

extract: <file> (html)
  title: Release notes
  2 part(s):
    h1 Overview | lines 1-2
    h2 Overview > Install | lines 3-4

--json adds the extracted text and each part's index, level, heading, trail, line_start, line_end and source_offset (the byte offset of its start tag in the source file), plus parts_omitted when more than 256 headings were found. A leading frontmatter block in the file is left out, as ingest leaves it out. Each heading and the title are redacted (key=value credentials and URL userinfo) over their first 4,096 code units before they are capped at 200 characters (a longer one ends in …), and in the parts list a backslash in a heading is written as \\ and a | as \|. The text has every <private> region replaced by its placeholder and is otherwise not redacted; it is a preview of a file the caller already reads, and it never reaches a page. For CSV and TSV it prints the delimiter and the counts, then the ## Table section the summary page would hold:

extract: <file> (csv)
  comma-delimited, 2 column(s), 2 row(s), 2 rendered, 0 redacted cell(s)
## Table

### Rows 1-2

```table
name | qty
bolt | 4
nut | 7
```

Any other file prints not extracted: <reason>, and --json returns {ok: true, path, extracted: false, format, reason, detail?}: format-read-verbatim for Markdown and plain text, format-not-extractable for a named format such as PDF, format-unknown (with format: null) for an extension the registry does not know (a named or unknown format is answered from the extension, without reading the file), and not-a-regular-file, source-too-large, not-utf8 or a table refusal for a file that could not be read or parsed. The exit code is 0 in every one of these cases: a file that cannot be extracted is a result, not an error. A path that does not exist or is a symbolic link is an error (exit code 1).

Knowledge packs

A knowledge pack is a selected subset of Brain knowledge - rules and the pages that hold runbooks and conventions - that another vault can preview, install as untrusted candidates, and remove as a unit. It is not a schema pack (the _brain.yaml schema: vocabulary block): the two never share a verb.

o2b brain knowledge-pack export    --name <name> --select <sel>[,<sel>...] --out <dir> [--version <v>] [--force] [--json]
o2b brain knowledge-pack preview   <pack-dir> [--json]
o2b brain knowledge-pack install   <pack-dir> [--agent <name>] [--json]
o2b brain knowledge-pack uninstall <name> [--confirm] [--json]
o2b brain knowledge-pack list      [--json]
  • Format. A pack directory is an OKF bundle of the selected pages (okf.json, concepts/…, no log.md), plus preferences.json (the selected rules as bank-export rows) and knowledge-pack.json (name, version, selection, a sha256 per file and a digest over all of them).
  • Selectors. pref-<slug>, a page id or vault path, topic:<topic>, tag:<tag>; comma-separated or repeated. A selector that matches nothing fails the export - a pack is never "whatever happened to match".
  • Privacy. Export BLOCKS a page declaring visibility:, any entry with an owner: claim, and an unreviewed OKF Review/ candidate, and names each on stdout. Everything carried goes through the shared egress redactor (registry entry brain-knowledge-pack-export) before the pack is sealed, so the hashes cover the redacted bytes; <private> regions are replaced. Rules leave without their evidence links and rendered body.
  • Preview shows the manifest, count, a guarded sample per entry, the integrity verdict, conflicts with the vault (id_exists, previously_retired, topic_claimed, review_target_exists, path_exists) and privacy / prompt-injection warnings. It exits 1 when the pack fails its own integrity check.
  • Install refuses a pack whose files, file set, name or version no longer match its digest. Rules land unconfirmed on a fresh trial window (dream.unconfirmed_window_days from the install instant) through the audited preference transaction, with the source vault's evidence, counters, revision, pin and aliases cleared; a rule whose id this vault already holds or once retired is skipped, never overwritten. Pages stage under OKF Review/ with okf_review: pending and machinery stripped. Every landed entry is stamped knowledge_pack: <name>@<digest12>, and each staged page also carries knowledge_pack_sha, a fingerprint of its body and authored frontmatter as installed; both are written by the installer and stripped from any bundle that supplies them itself.
  • Provenance of an installed entry shows in o2b brain query --preference (and brain_query) and in search trust metadata (brain_search with trust: true, field trust.knowledge_pack).
  • Uninstall is a dry run until --confirm. It removes every stamped rule and every page still staged under OKF Review/ for that pack name, behind a knowledge-pack-uninstall snapshot, and writes a source_invalidation continuity record for pack:<name>. A rule that gained evidence links in this vault, a page promoted out of the review lane, and a staged page whose body or authored frontmatter changed since install (reason edited) are kept and named. Staged pages live outside Brain/, so the snapshot does not cover them (the recoverability line says so); only untouched staged pages are removed, and re-installing the pack restores them.
  • MCP. CLI only: a preview reads an operator-named directory outside the vault, which no MCP tool does.

Stability and trust (since v1.0.0)

o2b brain dream               [run] [--dry-run] | stage | validate <run-id> | apply <run-id> | retriage <run-id> | discard <run-id> | list - staged lifecycle over a persisted proposal bundle; validate/apply exit 1 on drift
o2b brain dream retriage R    re-run the deterministic salience gate (`dream.salience_threshold`) against staged bundle R and report the delta: which facts would newly enter or leave the rollup ladder's fold set, each named with its current and its staged score. Read-only - the bundle is never rewritten, so re-stage to adopt the new partition. An unknown bundle, or one staged before the gate shipped, exits 2 naming the reason. Mirrors MCP `brain_dream` `action: "retriage"`
o2b brain dream --step S      run ONE independently-runnable step instead of the full pass: `scan` (pure read) or `heal-enrich` (asking for it is the opt-in, so the config gate is not consulted). Any other token - including a dream reporting phase - exits 2 with the specific reason that step cannot run alone and the runnable set. Returns a partial result marked `partial: true`, never a run summary. Mirrors MCP `brain_dream` `step`
o2b brain dream --gate N=V    override one phase gate for THIS RUN ONLY (`--gate heal_enrich=true|false`); repeatable, never written back to `Brain/_brain.yaml`, so a targeted pass needs no config edit and no revert. An unknown gate or a non-boolean value exits 2 naming the known gates. Mirrors MCP `brain_dream` `gates: {heal_enrich: bool}`
o2b brain dream (archive)     every run also moves inbox signals older than `dream.contradiction_window_days` that it did not consume into `Brain/inbox/archived/` (never deleted; `archived_signals` in the summary and the log event, previewed by `--dry-run`, count in MCP `archived_signals_count`). `dream.archive_stale_signals: false` turns it off; `o2b brain doctor` reports `inbox-archivable` while a pass is due
o2b brain doctor              gains the removed-tool-reference warning: vault notes, root instruction files, and installed skills naming a tool removed in 1.0.0 are flagged with the replacement
o2b brain doctor              opt-in `entity-alias-candidate` lint (off by default): with `entity_semantic_dedup_enabled: true` surfaces lexical entity-name variants ("Google LLC" vs "Google Inc") as PROPOSAL-ONLY alias-merge candidates via a deterministic jaccard layer (`entity_semantic_dedup_lexical_threshold`, default 0.8); never auto-merges or rewrites the identity key. The embedding-cosine layer (`entity_semantic_dedup_threshold`, default 0.92, reuses the configured embedding provider) is exposed as a library reader for apply plans
o2b brain daily | weekly | monthly | morning-brief | timeline
                              gain additive timezone + local_time JSON fields when `timezone:` is configured; storage stays canonical UTC
o2b brain digest | daily | weekly
                              with report_snapshots_enabled persist Brain/reports/<surface>/<date>.json and report a deterministic Since-last-run delta

Forward pointers (next: / next_command)

A verb that succeeds and leaves the caller with somewhere to go names that place through one mechanism. On a human stream it prints one next: <command> line; under --json the same command arrives as an additive next_command string field in the payload, and the key is absent whenever no exit resolves. The command is always a structural o2b invocation resolved from the diagnostics registry - never prose - so a caller can execute it verbatim.

Verbs that carry it today: o2b status, o2b brain init, bridges list|discover, clusters list|run, dream list, git status|mine, inbox-drain, intention list, intent-review, tune status, o2b search index|reindex|status. o2b brain tiers check deliberately carries none under --json: that state has two exits (restore or accept), the wire key is singular, and naming one would tell a machine caller the other does not exist.

When there is no command. About two thirds of the doctor's issue codes resolve to none, because the repair is a judgement over content or an edit whose target shape the finding cannot supply. Those are not silent: o2b brain doctor prints no exit: <code> - <reason> once per reported code, and --json (and MCP brain_doctor) carries the same reasons in an additive no_exit object beside the issue streams - once per code rather than repeated on each record, because a reason is about a class. Every doctor code is one or the other; tests/core/brain/doctor-exit-census.test.ts enumerates them from the detector and fails on a code that is neither, so an unregistered code can no longer mean "nobody got round to it". The key and the lines are absent whenever every reported code has an exit, and on a clean vault.

Long-running operations (dream, o2b search index | reindex | vector-backfill, bridges discover, clusters run, architect, the machine-wide session sweep, the maintenance lane) run under a cooperative safeguard deadline: safeguard_timeout_seconds (default 600, 0 disables, env OPEN_SECOND_BRAIN_SAFEGUARD_TIMEOUT) with per-operation overrides like safeguard_timeout_dream_seconds. The session sweep's key is safeguard_timeout_sessions_seconds: it is the one operation whose cost is dominated by bytes rather than by candidates - hashing transcripts that live on somebody else's disk - so its budget is worth setting separately from the passes that walk the vault. A tripped deadline aborts at the next checkpoint - between atomic writes - and reports {ok:false, timed_out:true} on exit 1; maintenance-lane task results carry timed_out per task. The frozen-surface policy lives in docs/stability.md; the 0.x to 1.0.0 migration table in docs/updating.md.

Watching one. The same verbs take --progress, which writes one newline-delimited JSON record per checkpoint to stderr - {"schema":"o2b.progress.v1","operation":…,"kind":…,"stage":…,"completed":N}, where kind is started, advanced, finished or stopped, and total is present only where a denominator is known before the loop starts (the embedding phase knows its pending count; the index walk consumes a generator and cannot). A CLI progress stream never carries refused - no counter emits it; it is an MCP-only shape on a different transport, so a CLI reader that handles it writes dead code. stage is an identifier from the operation's own phase vocabulary, never a sentence. Stdout is untouched, so a --json payload is byte-identical with and without the flag, and o2b search index --verbose keeps its separate per-file stream unchanged - the two answer different questions. Progress is opt-in for the same reason --verbose is; when a stream cannot carry it the verb says progress: not emitted (<reason>) rather than writing into a buffer nobody will read in time. o2b brain maintenance run --progress forwards the stream of whichever task is running rather than counting its own four steps, so every record names the operation that emitted it. o2b brain dream --step <s> --progress reports under dream with the step's own name as the stage (scan, heal-enrich) rather than the five stages of a full pass, and runs under the same dream deadline: a step asked for on its own IS the run. Under --json a tripped deadline comes back as {ok:false, step, timed_out:true, message}, the same marker the staged actions and the inline pass carry. The deadline is checked per file and per directory in scan, and per page in each of heal-enrich's two loops. Two phases of heal-enrich cross no boundary of their own and are therefore not interruptible at all: the vault listing, which walks every page and parses its frontmatter in one call, and the one-shot title/alias phrase build. The listing is the first thing the step does, so the only checkpoint it has is the one immediately AFTER it; the phrase build is bracketed on both sides, so an already-elapsed budget refuses to pay for it. Either way a step already inside one of them runs it to the end before it can stop. o2b search reindex --progress opens a lock stage before it waits for the writer lock (that wait can be the length of a competing rebuild) and does not report finished until after the swap stage that renames the staging build over the live index - the run is not over when the build is. o2b search vector-backfill --progress reports under reindex too, because the pass IS that run's embedding phase on its own: a plan stage while it counts, then the shared embed stage with the pending count as its denominator.

Which emitters those verbs actually have is not taken on trust. tests/cli/progress-emitter-census.test.ts enumerates every progressCounter( call site in src/ from the source, maps each to the entry point that reaches it, then RUNS each entry point against a fixture and requires records carrying that site's operation and stage to arrive with a terminator. A call site no entry point reaches is a failure, which is how vector-backfill - which had grown the whole spine in core with no flag to reach it - was found.

Stopping one, and what Ctrl-C actually does. This differs per verb, and the difference is a property of the operation rather than a gap in the wiring. A cooperative interrupt is delivered to a JavaScript signal handler, and a signal handler runs on the event loop; an operation that never yields to the event loop therefore cannot observe one, and merely registering a handler would suppress the default terminate and make the keystroke do nothing at all.

Verb Ctrl-C / SIGTERM
o2b search index, o2b search reindex Stops the run at the next checkpoint - between files, between embed batches, never mid-write - and exits 130 (SIGINT) or 143 (SIGTERM). A stopped rebuild leaves the live index exactly as it found it, because the staging build is abandoned before the swap.
o2b search vector-backfill Stops between embed batches and exits 130 / 143; vectors already written stay written, because each chunk commits as it is computed. The dry run and the planning query are synchronous SQLite between two awaits, so a keystroke landing there is not observed at a checkpoint - it ends the process on release instead, which is the same outcome the un-suppressed keystroke would have had.
o2b brain maintenance run Stops the lane at a task boundary and exits 130 / 143. The lane journals the stop and releases its lease, so the vault is never left leased.
o2b brain dream (including stage/validate/apply), o2b brain bridges discover, o2b brain clusters run, o2b brain architect, o2b brain import-session --status | --discover Terminates the process immediately, the ordinary shell behaviour. These passes are synchronous end to end, so no cooperative stop is possible and none is claimed. Every artifact they write is written atomically, so a killed pass leaves no half-written note - it leaves the vault as it was before the pass, or after the last completed write. Their deadline (safeguard_timeout_*_seconds, above) is the only cooperative stop they have.
o2b search watch Exits 0, unchanged: stopping is how that command ends.
o2b mcp (both transports, since v1.50.0) Stops accepting new requests, waits for the in-flight ones to a bounded deadline, closes, and exits 130 / 143. It is the one verb here that calls process.exit rather than re-raising, because two exit hooks - the search-store WAL checkpoint and the lock release - do not run when a process dies by signal. See "Shutdown and draining" in mcp.md.

o2b mcp exits 70 on an uncaught exception (since v1.65.0) after naming it on stderr: the served transports install a fault guard that survives, names, rate-limits and counts unhandled promise rejections, but an exception leaves the process state unknown, so the server exits - through process.exit, for the same two exit hooks - without draining. See "Background faults" in mcp.md.

For the two verbs that do hold a handle, a second interrupt is not intercepted and falls through to the default handler, so a wedged run is always killable by pressing the key twice. And an interrupt that arrives while such a verb is in a region with no checkpoint - opening a store, writing a report - is not swallowed: the verb prints interrupted: SIGINT arrived while … stopping now and ends with the signal's code rather than returning 0. A run the operator stopped never exits 0.

Lock and cache overrides (since v1.61.0)

Three environment variables tune how long a writer waits for a shared lock and where the machine-local dedup cache lives. None of them has a _brain.yaml key: they describe the host, not the vault.

Variable Default Effect
OPEN_SECOND_BRAIN_LOCK_WAIT_MS 5000 How long a write to shared ingest state (content manifest, plan checkpoint, session ledger, git record store) waits for its lock before it is refused with ELOCKED. A whole number of milliseconds; 0 means one attempt; anything else is an error. The one-second interactive wait is not affected. See Source pipeline integrity.
OPEN_SECOND_BRAIN_DEDUP_CACHE_DIR the user cache directory, open-second-brain/dedup-index/ Where the signal dedup index cache is kept, one <vault-digest>.json per vault, outside the vault.
OPEN_SECOND_BRAIN_DEDUP_CACHE on 0 turns the dedup index cache off; every capture then walks the inbox, processed/ and archived/ in full.

The automatic Brain upgrade worker takes no override: its lock is .open-second-brain/self-heal-upgrade.lock in the vault, so a hook, the MCP server and the worker agree on it whatever their temp directories are.

Vault scope

Single scope policy for every vault walker: vault.ignore_paths excludes, and the optional vault.include_paths allowlist narrows. A path is in scope when it is not excluded AND, if an allowlist is declared, under one of its roots. Absent, the allowlist changes nothing; an empty one is refused at parse time, because a list admitting no path is an off switch on indexing rather than a boundary. A dead include root is an error-severity vault-include-missing-path doctor finding — unlike a dead exclusion, it can leave the index empty.

Both keys share one grammar: a value with no slash is a bare name matched at any depth, a value with a slash is a vault-relative path matched exactly.

o2b vault status              Walks the vault under the active policy; reports include / exclude counts, the declared include roots, and which rule or polarity refused each excluded path
o2b vault inspect <relpath>   Point-check one vault-relative path; reports the scope verdict and which polarity refused it, the matched rule, source, whether the search index would admit the path, whether it exists on disk, and the write-binding verdict (none declared | admits | refuses)
o2b vault profile <sub>       Manage named multi-vault profiles (since v0.22.0): list | create <name> <vault> | switch <name>; pointer-based activation in profiles.json
o2b vault map [show]          Print the resolved vault-map role tokens -> folders (since v0.22.0), merging an optional Brain/_vault-map.yaml over defaults; read-only

Index admission joins the scope walk (since v1.46.0)

Scope is one question and index admission is another, and until v1.46.0 the two walkers disagreed: the search indexer refused lane-owned paths and o2b vault status did not, so status reported index coverage for files the indexer skips. The scope walker now applies the same admission filter the search walker has always applied, so the two answer alike.

This changes the counts o2b vault status prints for a vault that holds the overwrite-only operational-state lane at Brain/state/ - the one lane admission refuses today, because its rows are read directly and are deliberately never surfaced through FTS, vector, or graph recall. The filter applies whether or not the vault declares vault.include_paths; a vault with no Brain/state/ directory counts exactly as before.

A refused lane path appears under excluded with the reason not-admitted, beside the two scope reasons that already existed. The reason field on each excluded entry is that closed set:

reason Refused by
ignored a vault.ignore_paths rule
not-included an allowlist is declared and the path is under none of its roots
not-admitted index admission, with no scope rule involved - rule and kind are null

The walk records one entry per refused subtree root rather than one per descendant, so Brain/state appears once and its files are not enumerated separately. That is the same rule an ignore exclusion already followed.

o2b vault inspect <relpath> reports the two verdicts SEPARATELY, because a path can be perfectly in scope and still never be indexed. --json carries status plus reason for the scope verdict (ignored, not-included, or null) and index_admitted plus index_refusal for the admission verdict (exact-state-lane, or null when the path is admitted). The human transcript prints an index: not indexed (<reason>) line under an included path whose admission was refused, and prints nothing there when the path is admitted.

State surfaces (since v1.50.0)

o2b state status              One row per declared state surface: where it resolved, whether it is reachable, which layer put it there
o2b state migrate --to <dir>  Plan the move of every in-vault state surface to <dir>; --apply --yes performs it, --dry-run is the explicit form of the plan
o2b state rollback --from <dir>
                              Put back what the migration's manifest binds and whose digest still matches; --to <vault> names a vault that has moved since

All three take --json; status and migrate also take --vault and --config.

Three questions an operator asks before they copy, migrate or delete a vault had no surface that answered them: where each surface lives on this machine, whether it is reachable, and which layer put it there. o2b state status is that surface, and the vault_health MCP tool carries the same value under state_surfaces - one report rendered two ways rather than two hand-written copies.

The catalogue and the two tiers

45 declared surfaces, printed whole. A row says what this build CAN keep at that location, never that this machine has it; presence is the measured half. They are grouped by what losing one costs, which is the first thing a migration needs:

Tier Meaning
derived rebuildable from the vault, so deleting it costs time and not memory
vault-content the memory itself, so a loss here is permanent without a backup

The tier is a RECOVERY story, not a location: Brain/.state/anticipatory/ sits inside the Markdown tree and is derived because deleting it costs one recomputation, and the search index sits outside it and is derived for the same reason. Whether a surface can hold memory CONTENT is a separate axis, reported per row and counted in the summary line - 17 of the 45 can.

Only two overrides move anything: OPEN_SECOND_BRAIN_SEARCH_DB / search_db_path relocate the search store and everything that follows it, and O2B_DEVICE_ID / device_id select which shard of the log, the truth ledger and every other per-device ledger this machine writes. The second is reported as provenance and not substituted into the path, because it names a file inside a directory the inventory reports at directory granularity. Every other row prints at its only location - nothing can move it.

Reachability is a tri-state, and it carries its reason

An unreadable surface is not an absent one. A boolean would have to spell "I could not look" as false, which reads as "it is not there" at every call site and sends an operator to recreate something that already exists under a mode they cannot read.

state Rendered as Meaning
present present the probe found something at the resolved path
absent absent - nothing has created it in this vault yet absence is a state, not a failure
unchecked NOT CHECKED - could not be probed (<errno>): <message> an EACCES on the parent, an ELOOP, an EIO from a dying disk

o2b state status exits 0 whenever the inventory COMPLETED, including over an unchecked row: turning a reporting verb into a gate would duplicate o2b doctor. The probe falls back to lstat after stat, so a surface whose root is a dangling symlink answers present rather than absent - something IS at that path, and saying so is what carries it to the migration check that knows how to explain a link.

Migration refuses before it commits, and returns every reason at once

o2b state migrate runs every check and returns ALL of them; it never short-circuits on the first refusal and never touches the filesystem while planning. An operator who fixes a symlink only to be told about a held writer lock, then about a full disk, learns their vault's problems one interrupted migration at a time.

Seven refusal classes, each a distinct repair:

Code Refused because
symlink a bound path is a symbolic link. A migration copies bytes, so the link would arrive pointing at the OLD location and its target would silently stay behind
special_file a FIFO, socket or device node, which has no bytes to copy; copying it would produce a plain file that only looks like the original
destination_occupied the destination exists and is not an empty directory, or could not be examined or listed. Merging would make the manifest unable to say which files this migration put there
insufficient_space the destination filesystem has less free space than the plan needs, or its free space could not be measured at all
reserved_namespace the destination is inside the vault, an ancestor of it, the filesystem root, or the home directory itself. Every comparison is made over the REALPATH, so a symlink cannot smuggle the state tree back into the vault it is leaving
writer_lock_held the search index writer lock is held by a live writer, or its state could not be determined. Bytes appended after this run hashed the index would land in neither copy - the one loss a digest cannot detect afterwards
unreadable_surface a declared surface, a directory in the tree, or a file in it could not be probed, listed or read. A tree this run cannot enumerate cannot be bound by a manifest, and an unbound file is not restorable

Every refusal names what was found AND the remedy. A surface an override put outside the vault is not a refusal: it is reported separately as left where it is, because migrating the vault does not move it.

The dry-run ladder is o2b brain upgrade's, verb for verb, copied rather than reinvented because an operator who has learned one destructive verb in this CLI has learned all of them:

Form Behaviour
bare --to <dir> plans and prints; writes nothing
--dry-run the same plan, said explicitly. Mutually exclusive with --apply
--apply --yes performs it
--apply on a TTY prints the plan, states how many files and bytes are about to MOVE and where the rollback manifest will be, and prompts; only y or yes proceeds
--apply under --json or a non-TTY stdin refused - nobody can answer the prompt

Within the apply the same rule holds one level down. Every file is COPIED and its landed bytes re-digested BEFORE any source byte is removed. A copy that fails half way unwinds what it wrote - including the empty directory skeleton, so the identical retry is not blocked by the abort - and leaves the source exactly as it found it. Only once the whole tree has landed and verified does the removal pass run, and it prunes only directories that are now empty, so an operator's own file inside a state root keeps its parent alive.

The destination gets state-migration.json: a schema version, the digest algorithm, the source type, the absolute source and destination roots, the minimal set of vault-relative directories covering every moved file (canonical_roots - a declared surface nested inside another is absorbed by the one that contains it, so no byte is bound twice), a byte count and a SHA-256 per file, and one digest over all of it.

Rollback needs --from, and takes --to

--from <dir> is the directory a migration was moved TO - the directory holding its manifest - and it is required. --to <vault> is optional and answers a vault that has MOVED since: the manifest records an absolute source root, and a rollback into a dead path recreated the directory, copied into it, removed the destination copies and exited 0, leaving the live vault with nothing. A recorded root that is no longer a directory is therefore refused BY NAME before a single entry is classified, and the plan prints the recorded root beside the override either way, so restoring somewhere else is something the operator reads rather than remembers typing.

The ladder is the migration's: a bare invocation (or --dry-run) plans and prints, --apply --yes performs it, and --apply without --yes prompts on a TTY and is refused under --json or a non-TTY stdin.

Only what the manifest binds AND whose digest still matches goes back. Everything else is refused by name and left exactly where it is: nothing this verb can do makes an operator's later edit recoverable, so the one thing it must never do is delete it.

Code Left alone because
digest_mismatch the destination copy has changed since the migration
missing_at_destination the manifest binds it and nothing is there now
source_diverged something has written the source path since the migration
unreadable this run could not read one end of it, so it cannot be shown to be the file the manifest binds

Every entry is measured AGAIN at apply time, because a plan is a statement about the moment it was made. The manifest is removed only when there is nothing left to refuse from either half - an incomplete rollback keeps it, since the refused entries are the only record of where those bytes came from.

A manifest that does not verify against its own digest, declares a schema version this build does not read, or is not parseable JSON is refused rather than repaired: a rollback driven by an edited manifest is a restore of something nobody measured.

Exit codes

Code Meaning
0 the inventory completed, the plan was printed with no refusals, or the operation did what it was asked
1 a REFUSAL: a migration that will not run, a rollback that left files behind, or a plan that could not be built
2 a usage error: a missing --to or --from, --dry-run with --apply, an unknown subcommand, or --apply without --yes where nobody can answer a prompt

The split is the point. A refusal is an ANSWER - the command did what it was asked and the answer is no - and a supervisor that cannot tell it from a broken invocation retries the wrong one. A rollback that restored four of five files exits 1 even though everything it DID do succeeded, because "restored 4 of 5" is not a success.

Discipline (daily logging cron)

o2b discipline report         Render the daily MarkdownV2 block to stdout (brain-event counts per agent vs git/mtime/vault activity plus complexity-to-thinking ratio); status ok | info | alert
o2b discipline install        Register the Hermes cron job that delivers the report. --telegram-target is required; --at defaults to "59 4 * * *" UTC; --weekly installs a Monday 08:59 weekly digest. The jobs file is `~/.hermes/cron/jobs.json` of the running user (since v1.65.0; before, always under the root user's home); `OSB_HERMES_JOBS` overrides it
o2b discipline uninstall      Remove the cron job; --weekly removes only the weekly digest, without flag removes both

See hermes-cron.md for the cron envelope and Telegram delivery shape.

Partner (read-only, since v1.12.0)

Reports on external code-project partners. Strictly read-only: never installs, initializes, extracts, or mutates a partner index or the vault.

o2b partner codegraph report  Resolve the in-scope code project and report the codegraph index state (no_project | absent | not_indexed | indexed with node/file/edge counts | error) plus a structural Cargo.toml workspace-member list. When indexed, runs a read-only, non-blocking graph-health gate (index.health) that flags empty-graph, collapsed-edges, dangling-references, self-loops, and cache-root-mismatch before labeling/import/recall trust the graph. Non-Rust projects report cargo_workspace: null with a reason. --vault sharpens the scan scope; --json emits the schema-versioned report; --fail-on-health exits 1 unless the index is present AND its health gate is clean, so a scheduled job can branch on it (without the flag the exit stays 0 whatever the report says)
o2b partner codegraph resync  Print a cron recipe that re-indexes the in-scope code project whenever its commit moves. --cron-template is REQUIRED (there is nothing else this verb does, and it never runs an indexer); --interval accepts <N>m|h|d and defaults to 6h; --format cron|systemd (default cron, since v1.65.0) prints the crontab line or a systemd user timer pair; --project names the repository instead of resolving it from the scan; --vault sharpens that scan. Output is text on stdout and nothing else: the emitted script aborts on a cache-root-mismatch health warning, aborts when jq is missing rather than matching the report loosely, skips quietly when the commit is unchanged, invokes codegraph itself, and records the commit in a stamp under ${XDG_STATE_HOME:-$HOME/.local/state}/open-second-brain only after a --fail-on-health check passes

Every write in the resync recipe is a shell command the operator's own crontab runs on the operator's own host. Open Second Brain never installs, initializes, or writes data for codegraph, and rendering the recipe creates nothing - which is also why there is no maintenance-lane task that re-indexes on its own.

Search

o2b search "<query>"          Hybrid full-text + semantic search across the vault
                              --property type=decision --property status=open
                              filters on frontmatter scalars (post-FTS phase)
                              --degree backlinks=0 / --degree outlinks>=5 filter by graph
                              cardinality (orphans, hubs; ops = != > >= < <=, ANDed)
                              --query-doc '<lanes>' separates intent/lex/vec/hyde recall lanes
                              --evidence-pack adds matched/missing term diagnostics, abstention text,
                              IDF-weighted coverage, per-token union records, and a completeness verdict
                              --since/--until scope recall by event time - frontmatter validity,
                              else the body-derived anchor, else mtime (ISO date/datetime,
                              today / yesterday / last week / last month, or 24h / 7d / 2w shorthand)
                              --include-superseded keeps superseded predecessors undemoted (history mode)
                              --verbose adds per-result why_retrieved reasons
                              --explain appends the retrieval decision trace and the memory trust
                              assessment (what the trust gate evaluated, surfaced and excluded, with
                              reasons); with the gate off it says so and names the switch that enables it
                              --json for structured output (includes reasons[])
                              --json carries retrieval_trail when the answer narrowed or came back
                              empty; the human transcript names the cause on the no-results line
                              total is the pre-truncation ranked pool, not the number of rows returned
                              CJK text is expanded for FTS recall without polluting returned content
                              --profile fast|balanced|thorough picks a recall preset; no profile leaves ranking unchanged
                              --disclosure cards returns compact cards (path, title, score, snippet, line range) instead of full content; default full
                              --match-mode all|any picks the keyword match breadth: all (default) requires every
                              term, any matches a note carrying any one term; typed AND/OR/NOT/NEAR stay dropped
o2b search expand             Drill one card: --chunk <id> (printed on the card) returns the fuller note and the raw chunk transcript, paged by --cursor; read-only
o2b search feedback           Record explicit recall feedback for one result
                              (--query Q --result <path> --verdict up|down; one JSON event file
                              under Brain/search/feedback/, learned weights refresh deterministically)
o2b search weights            Show base weights, learned multipliers, event count, and bounds
                              --reset removes the derived learned-weights file (events kept)
o2b search focus set          Persist a 120-minute ranking focus (--query Q and/or --path P; --ttl-minutes N; --session S binds it to one session, since v0.37.0)
o2b search focus status       Show the active focus; --json emits { active, focus }
o2b search focus clear        Clear the persisted focus file next to the search index (--session S clears one session's focus)
o2b search reindex            Rebuild the SQLite + FTS5 index from scratch
                              (required after upgrading to v0.13.0 recall schema or v0.26.0 CJK FTS content)
                              --force-cost bypasses the embedding cost gate for this run (since v0.36.0)
                              --embeddings computes vectors; --concurrency N, --db PATH, --verbose
                              --cron-template prints a periodic-reindex cron recipe on stdout and
                              writes nothing at all (no index run, no file, no scheduler entry)
                              --interval <N>m|h|d sets that recipe's cadence, default 30m; an interval
                              cron cannot express (seconds, 60m+, 24h+, 28d+) is refused with the
                              reason rather than rendered as a schedule that means something else
                              --format cron|systemd selects the recipe's scheduler (default cron;
                              systemd prints a user .service/.timer pair, since v1.65.0);
                              --interval or --format without --cron-template exits 2 and runs no reindex
                              --self-heal <run-id> records this run's terminal outcome (or its
                              failure, by name) on the self_heal_reindex metrics surface under that
                              id; set by the detached post-upgrade rebuild, whose streams are all
                              ignored, and whose parent mints the id so the two rows pair without a
                              machine-local pid
o2b search index              Incrementally update the index; --embeddings computes vectors, --progress watches it
                              --force-cost bypasses the embedding cost gate (since v0.36.0)
                              --freshen <token> [--freshen-state <dir>] is the background run freshen on
                              read starts (since v1.76.0): lowered priority, no output, outcome in
                              <dir>/freshen-state.json; meeting another index writer is a skip
o2b search vector-backfill    Run the vector phase ALONE for indexed chunks that have no vector -
                              no vault walk, no re-chunking, no frontmatter pass (since v1.43.0).
                              Dry-run by DEFAULT: it counts the pending chunks, reports the
                              configured semantic capability tier, and contacts no provider.
                              --apply is the only path that reaches a provider or writes a vector
                              --force-cost bypasses the embedding cost gate for that run
                              --path <prefix> scopes the run to chunks under a vault-relative
                              prefix; repeatable, and validated like every other path prefix (an
                              unsafe prefix is refused by name). The pending census, the estimate,
                              the cost gate and the spend receipt all read the same scoped census,
                              so `--path Brain/preferences/` prices and embeds only belief notes.
                              The text report adds a `scope:` line and the next step it names keeps
                              the scope. On Windows a prefix that would need quoting is not
                              spliced into the advice: next_command is omitted and the text says
                              `next: rerun this command with --apply and the same --path flags`.
                              The prefix is a raw string prefix, not a directory: end a directory
                              with `/`, or `Brain/pref` also matches `Brain/preferences-old/`. A
                              leading `./` is dropped and `\` becomes `/`; an empty prefix is
                              refused with INVALID_INPUT. A prefix that matches no indexed
                              document is warned on stderr by name (`scope <prefix> matches no
                              indexed document`) rather than reported like a fully embedded
                              scope (since v1.72.0)
                              --progress watches it; Ctrl-C stops it between embed batches
                              --json emits dry_run, capability_tier, capability_code, chunks_total,
                              pending, embedded, retries, estimated_cost_usd, price_source, and
                              path_prefixes on a scoped run, plus unmatched_path_prefixes when a
                              prefix matches no document. Since v1.72.0 estimated_cost_usd is
                              null (never 0, never omitted) when the model's price is unknown, and
                              the text report prints `price unknown` for it. Since v1.72.0, when
                              the configured gate would refuse the run unforced, --json adds
                              gate_blocked: true and gate_reason (unpriced or over_cap), the
                              dry-run text report adds `cost gate: would refuse (<reason>); add
                              --force-cost or ...` naming the price pair or embedding_cost_gate_usd;
                              the next step it names stays unforced. An --apply run that
                              reached the provider adds spend {model, tokens, estimated_usd,
                              price_source, forced}, its receipt. Both are absent otherwise
                              Idempotent; an --apply run that wrote vectors appends one
                              vector-backfill Brain log event
o2b search status             Index status; since v0.36.0 also reports the active embedding
                              signature (<provider>:<model>:<dimension>) and a refresh-cost estimate
                              Since v1.72.0 the estimate is null and printed as `price unknown` when
                              the model has no known price; --json adds refresh_price_source
                              (builtin, operator or unknown)
                              Since v1.64.0, once an index exists, status also prints
                              event_time: <with>/<documents> documents (earliest <ISO>, latest <ISO>,
                              <n> in the last 30 days); --json carries it as the event_time object
                              (documents, with_event_time, earliest, latest, recent_window_days,
                              in_recent_window)
                              Since v1.76.0 it prints freshen (every <n>s or off), index_age, last_freshen
                              (completed <ts> (<n> changed) / failed <ts> (<n> in a row): <error> /
                              (none)) and freshen_backoff (until <ts>) while a backoff is active; --json
                              carries them as the freshen object (interval_s, index_age_s, last_outcome,
                              last_run_at, last_changed, last_error, failures, backoff_until)
o2b search check              Pre-flight diagnostics: vault, index directory, SQLite/FTS5, the
                              vector extension, the embedding key, the provider, the vector ABI stamp
                              --integrity additionally runs a full PRAGMA quick_check over the index
                              file, on demand and regardless of the store's 24-hour interval gate.
                              Opt-in because the scan is linear in the index size (~22 ms per MB, so
                              about 30 s on a 1.3 GB index); the size and a time estimate are printed
                              on stderr before it starts and the time spent is reported after.
                              The verdict lands in the same integrity_checked_at / integrity_fault
                              cells the write-open gate writes, so a condemned file is refused by
                              every later read; a fault names o2b search reindex and exits 1.
                              A store that has never been scanned reports previous_full_check: never -
                              the absence of a fault is never reported as a pass.
                              --json adds an integrity object; without the flag the output is
                              unchanged and no index_state cell is touched
                              --no-probe skips the live embedding-provider call. The probe is on by
                              default, as it has been in every release that resolved a key; the flag
                              exists so a diagnostic can be run with no network at all.

                              Exit codes. `0` clean. `1` a machine fault or a condemned index, and it
                              keeps precedence: a provider verdict never masks one. `5` the embedding
                              provider is configured and was proved unreachable - it answered, and it
                              refused. `6` the probe did not complete, by its own clock or by the
                              outer budget; "I could not find out" is not "it is broken", and the two
                              are separate codes for that reason.

                              Code `5` is a behaviour change, and a breaking one: this verb exited `0`
                              over a provider it had already proved unreachable, so a script gating on
                              the exit code read it as healthy. It is deliberately the same number
                              `o2b install --check` uses for the same condition, and a test asserts the
                              two cannot drift. A provider that is not configured still exits `0` -
                              absent is not broken.

                              The `--json` key `provider_reachable` (a boolean, or null when unknown)
                              is replaced by `provider_probe`, a string from a closed vocabulary:
                              not-configured, reachable, unreachable, timed-out, skipped. The boolean
                              had to answer four questions with two values, so a provider that refused
                              and one that never answered were the same `false`.

                              Two censuses ride the report in EVERY state, because a check that could
                              not measure something must not answer with silence. `pending_vectors` is
                              the measured count of chunks with no vector - `{verdict: measured,
                              pending, chunks}`, or `{verdict: unrecorded, reason}` when the index is
                              absent or will not open, which is never reported as a count of zero. It
                              is what the reindex recommendation now gates on: a fully embedded vault
                              is told nothing, an index with no vectors at all is told to compute its
                              first ones, and a partially embedded index is pointed at
                              `o2b search vector-backfill`, whose dry run prices the work.
                              `embedder_record` is the record-vs-data audit: the dimension
                              `index_state` claims against the widths the `embeddings` rows carry and
                              the width `chunk_vec` declares. Its outcome is `complete` or
                              `contradicted` - a record the data itself disproves, which is a
                              different finding from ABI drift (the record disagrees with this build)
                              and from an unrecorded token (no claim was ever made).

                              Since v1.72.0, when no embedding key resolves, the report names where
                              it looked, by name only: `key_sources_checked:` lists the sources in
                              probe order (OPEN_SECOND_BRAIN_EMBEDDING_KEY, embedding_api_key, then
                              the env-key names of the registered profile `embedding_provider`
                              selects), and `key_present_under:` lists the other registered profiles
                              whose env key is set, or says `none`. --json carries
                              the same as `credential_sources` {consulted, present_elsewhere}. Only
                              names you declared are consulted (config keys and your provider
                              registry), no value is ever printed, and a configured setup's output
                              is unchanged. The recommendations also name an embedding model with
                              no known price (with the price pair to declare and what a positive
                              gate does with it) and a price pair that names a model other than the
                              active one.
o2b search restamp            Record this build's sqlite-vec version as the one the stored vectors are
                              accepted under - the repair for a drift confined to
                              embedding_vec_version, which is an ABI marker rather than a property of
                              any vector (two peers on different sqlite-vec builds each read the
                              other's index as drifted). Dry-run by DEFAULT: it prints the change it
                              would make and writes nothing. --apply writes that one index_state cell
                              and nothing else. It REFUSES by name when the recorded model or
                              dimension disagrees with this build - those describe the vectors
                              themselves, and repairing them needs re-embedding under a verified
                              identity, which is deferred by design; the refusal names
                              `o2b search reindex --embeddings` instead. No path contacts a provider.
                              --json emits dry_run, field, recorded, runtime, changed, applied
o2b search provider add NAME  Register an OpenAI-compatible embedding endpoint (since v0.36.0)
                              --base-url U --model M --env-key K (K is the env var NAME holding the key);
                              persisted to Brain/search/embedding-providers.json, resolved after built-ins
o2b search provider list      List registered provider profiles (--json for the array)
o2b search provider show NAME Show one registered profile (--json)
o2b search provider remove    Remove a registered profile by NAME
o2b search rerank-fit         Per-store reranker fit check (read-only diagnostic): samples real
                              recorded queries and correlates the reranker's scores with the base
                              retrieval signal. Reports fits (quiet), out_of_domain (low fit), or
                              inverted (negative) with a disable/swap recommendation; a rerankerless
                              vault reports inapplicable. --max-queries N --top-k K --json
o2b search rerank-eval        Rerank eval gate over a labelled dataset: runs the recall benchmark
                              with rerank off and with --kind local|decision-model|openai-compat
                              (default local) and reports hit@1/3/5/10 (up to --k), hit@k and MRR
                              for both, the deltas, per-query wins/losses/ties against off and the
                              recommendation (enable only on lift without a hit@k regression); for
                              decision-model also the run's calls by outcome, cost and latency.
                              --dataset PATH (required) --k N --compare-local --json. The
                              decision-model arm runs in enforce and needs an active decision-model
                              config; every query then sends its top candidates to that endpoint.

search_rerank_kind is openai-compat (default), local or decision-model. The decision-model kind reranks through the optional decision model and its decision_model_* config; see docs/decision-models.md.

o2b brain doctor checks the rerank configuration (since v1.66.0) whenever search_rerank_enabled is on with the openai-compat kind. The check reads configuration only and sends no request:

Code Stream When Next command
rerank-endpoint-unconfigured error the base URL is missing, carries user:password@ credentials or is not an accepted endpoint (the configured value is never repeated in the finding; <search_rerank_base_url> stands in for it), the model or the key is missing or blank, or search_rerank_provider names no registered profile and so left one of them empty, so every rerank-enabled search would fail o2b search rerank-provider list
rerank-model-sunset-announced warning the configured model has an announced decommission date that is 90 days away or fewer, or already past o2b search rerank-provider add
rerank-model-sunset-unsurveyed uncertain the configured model is outside the shipped rerank decommission survey, so no statement was made about it none, with the reason printed
rerank-model-sunset-undetermined uncertain the check ran and reached no verdict, for example because the survey is older than its horizon none, with the reason printed

The survey records model strings and the published notice each entry rests on, never an endpoint: a hosted service shutting down while its open checkpoints keep running elsewhere is not a model sunset, and is reported at query time as rerank-provider-unavailable instead.

The retrieval trail (since v1.46.0)

A search that came back with nothing used to say only (no results), which is the same sentence for an exhausted corpus and for an embedding provider that could not answer. o2b search --json and the MCP brain_search tool now carry a retrieval_trail key recording what the retrieval actually did:

  • retrieved - rows handed back, whether they rode results or cards. This is not total, which is the size of the ranked pool the window was cut from.
  • pool - that ranked pool size, the same number total reports.
  • degraded - the narrowings, in pipeline order. Each entry is a code plus an optional detail object carrying identifiers and integers only, never a provider message, a filesystem path, or your query text.
  • empty - present only on a zero-result answer that no degradation accounts for, carrying state, reason, and unknown_reason when the state is unknown. When a lane degraded, that degradation IS the explanation, and claiming the corpus holds nothing would be a stronger statement than the search can support.

The key is absent on a healthy answer - rows came back and nothing narrowed them - so an existing consumer's payload is unchanged.

degraded[].code is a CLOSED vocabulary. Every member is a stable identifier safe to branch on, and each names its own lane in its own name, so there is no separate lane field that could drift from it:

Code The narrowing it reports
index-stale the index was last updated more than ten minutes ago (detail.ageSeconds), so notes changed since then may be missing; the code reports the index age only and does not say a refresh is running: freshen on read starts a background run when it can, and starts none when it is off, a failed run is backing off, another run or writer holds the index, the index is read-only for this reader, or the run cannot be spawned
keyword-fts-match-empty the query tokenised to an empty FTS match, so the keyword lane never ran
keyword-trigram-lane-fault the trigram candidate lane could not be read; detail.fault carries that lane's own classification
semantic-embeddings-absent the index holds no compatible embedding
semantic-vec-extension-unavailable sqlite-vec is not loaded on this machine
semantic-capability-blocked the configured semantic capability blocks the vector lane; detail.tier names the rung
semantic-cost-unpriced the embedding model has no known price and embedding_cost_gate_usd is positive, so the query embed of a caller that is not local was refused before any provider call and the semantic lane did not run
semantic-query-truncated the query was longer than the effective embedding input window, so the semantic lane searched a cut prefix of it; detail.windowTokens is the window
semantic-query-empty-fit the instruction prefix alone fills the effective embedding input window, so no part of the query was left to embed and the semantic lane did not run; detail.windowTokens is the window
semantic-provider-unavailable the embedding provider could not answer; detail.category carries the error category
semantic-empty-query-vector the provider answered with an empty query vector
semantic-structured-lanes-skipped a structured semantic lane was requested while semantic search is off
semantic-embedding-abi-drift the index carries embeddings written by another build; detail.fields counts the contradicted ABI fields
hybrid-degraded hybrid recall was asked for and the semantic lane did not run, so this answer is keyword-only
hybrid-deadline-exceeded the composite hybrid path (embed, semantic top-k, rerank, second pass) outlived search_hybrid_deadline_ms, so the phases past the budget were cut; detail.budgetMs is the deadline, detail.elapsedMs the moment it fired
rank-cap-truncated-pool the rank cap truncated the candidate pool; detail.cap is the cap that bit
rerank-provider-unavailable the configured cross-encoder reranker (remote endpoint or local model) could not answer, so the answer keeps the heuristic order; detail.category is one of auth, quota, gone, rejected, transient, timeout, network, malformed, unclassified
rerank-model-sunset the configured rerank model's announced decommission date has passed, so no rerank request was sent and the answer keeps the heuristic order
relevance-floor-dropped-rows the relevance floor dropped ranked rows; detail.dropped counts them
scope-filters-dropped-rows visibility, ownership, or session / project scope dropped ranked rows; detail.dropped against detail.before
cross-vault-origin-failed a cross-vault origin could not be searched; detail.origin is the origin label
cross-vault-chain-stopped an origin answered confidently and the remaining origins were deliberately not searched; detail.skipped counts them

Members are not invented for conditions nothing reports, so every code above has a producer on the search path today. hybrid-degraded is the umbrella over the five semantic-* codes: they say why the lane did not run, it says what the caller received.

The two rerank-* codes report the optional cross-encoder rerank (since v1.66.0). Neither turns a search into an error: the answer keeps the heuristic order, and the existing rerank_degraded: warning stays the human signal for a failed endpoint. detail.category on rerank-provider-unavailable is computed from the typed failure, never from the provider's message:

Category The failure it names
auth the endpoint answered 401 or 403
quota the endpoint answered 402
gone the endpoint answered 404 or 410, the usual sign of a retired endpoint
rejected the endpoint answered any other 4xx
transient the endpoint answered 408, 429 or a 5xx. Some vendors also answer 429 for exhausted quota; the category is computed from the status alone
timeout the request outlived its timeout, including a 2xx body that stalled after the headers arrived
network the request never reached a complete answer: refused connection, DNS, TLS, a redirect, or a body that broke while it was read
malformed the endpoint answered, but the body was not JSON, carried the wrong number of scores, or an out-of-range or duplicate index
unclassified the rerank provider threw an error the cross-encoder did not type; it is named rather than folded into another category

An answer carrying rerank-provider-unavailable is served but never written to the query cache, so the next identical query asks the endpoint again instead of replaying the failure. An answer carrying rerank-model-sunset depends only on the build and the date, and is cached as usual; the cache key carries the survey's review date, so a build with a corrected survey does not replay the old answer. The sunset skip applies to the openai-compat kind only: the local and decision-model kinds carry no model string to look up.

The human transcript names the cause instead of printing a bare no-results line. The first degradation wins, because the lanes push in pipeline order and the earliest narrowing produced the ones after it:

(no results: the sqlite-vec extension is not loaded, so the semantic lane could not run)

With nothing degraded, the corpus statement answers instead - and that is a claim about this vault's index, not about your query. state is not_found, unknown, or did_not_happen; an unknown state prints its named reason in brackets (index-absent, index-stale, coverage-divergent, coverage-unavailable, index-instant-unusable, embeddings-incomplete):

(no results: unknown [index-stale] - <the verdict's own reason>)

With neither a degradation nor a corpus statement, the line is exactly what it always was. The English sentences belong to the transcript alone; the machine surfaces carry codes.

Embedding providers (since v0.36.0): embedding_provider accepts the built-in openai-compat, the offline local feature-hashing embedder (no cloud, no key, no model download; embedding_dimension default 256), disabled, or any name registered via o2b search provider add. embedding_cost_gate_usd (default 0 = off) refuses an embedding run whose estimated spend exceeds it unless --force-cost. Since v1.72.0 a blank (whitespace-only) value of the gate or of its env twin OPEN_SECOND_BRAIN_EMBEDDING_COST_GATE fails config resolution with INVALID_INPUT instead of reading as a gate of 0, and a blank search_rerank_min_score is refused the same way.

Embedding prices (since v1.72.0). Every estimate names where its price came from: builtin (the frozen price table, and the local embedder, which is free), operator (declared by you) or unknown. Declare a price for a model the table does not list, or correct a table price, with the operator price pair:

embedding_price_model: nomic-embed-text:latest
embedding_price_usd_per_mtok: 0.02

The env twins are OPEN_SECOND_BRAIN_EMBEDDING_PRICE_MODEL and OPEN_SECOND_BRAIN_EMBEDDING_PRICE_USD_PER_MTOK, and they win over the config keys as a pair: when either env twin is set, both halves come from env and the config pair is ignored, so an env model never pairs with a config rate. Set both keys or neither (a blank value counts as unset); the rate is USD per million tokens, a plain non-negative decimal number no larger than 1000000, and 0 declares the model free. A half pair, a negative rate, a non-decimal rate (0x10, 1e3, Infinity) or a rate above 1000000 fails config resolution with INVALID_INPUT naming the key or env variable that supplied it. The pair binds the price to one model name (compared case-insensitively), so switching models never re-targets it silently: o2b search check flags a declaration that names a model other than the active one. A price is not part of the embedding identity, so declaring or editing it never triggers a reindex. A loopback embedding_base_url is not assumed to be free; declare 0 for a local server that costs nothing.

An unknown price is reported as unknown, never as $0: the maintenance banner, the backfill dry run and search status print price unknown, the JSON estimates are null, and spend receipts carry price_source. A fully embedded index has nothing pending to pay for, so search status then prints no refresh_cost_est line and its JSON estimate is 0. Under a positive embedding_cost_gate_usd, an embedding run on a model with no known price and pending chunks is refused with EMBEDDING_COST_UNPRICED, because an unknown price cannot be checked against a cap. The message names the model, both price keys and --force-cost; the provider is never contacted. --force-cost passes the refusal, and the receipt then records forced: true, price_source: unknown and a null estimate. With the gate at 0 (the default) nothing is refused.

Query embeds (since v1.73.0). Every paid query embed passes one gate: the search lane (reached by o2b search, brain_search, brain_recall_feedback, brain_file_context, brain_eval, brain_benchmark, brain_tune and the recall-inject and gap-promote hooks) and the brain_context_pack semantic belief order. Under a positive embedding_cost_gate_usd, a caller that is not local is refused the query embed of a model with no known price before any provider is called. A search that asked for the semantic lane by name fails with EMBEDDING_COST_UNPRICED; a hybrid search falls back to keyword-only, says so in a warning and records semantic-cost-unpriced in its trail. The message names the model and the price pair that clears the refusal, and says only that the gate is positive, never its amount. The CLI runs at local reach, so o2b search is not gated, as before. The hooks run at remote reach: under a positive gate on an unpriced model they recall by keyword only and name the code on their local audit line (retrieval_degraded). Declare the price (0 for a free self-hosted model) to bring the semantic lane back.

The same gate fits the query to the model's input window before it is sent. The effective window is embedding_input_window_tokens when set, then the window the curated model table declares, then unknown; an unknown window cuts nothing. The cut counts the instruction prefix the provider sends, is made at a code-point boundary under the conservative token estimate, and is disclosed: a warning names the window and how much of the query was embedded, and the trail records semantic-query-truncated with detail.windowTokens. When the instruction prefix alone fills the window nothing is embedded: an explicit semantic search and the semantic belief order refuse with INVALID_INPUT, and a hybrid search falls back to keyword-only with the trail code semantic-query-empty-fit, which, unlike a cut, always means the semantic lane did not run. An answer refused by the gate is never cached, so declaring a price takes effect on the next search. An answer cut to the window is cached under a key that carries the effective window and the query prefix, so a repeated long query is not embedded again, and a declared or changed window takes effect on the next search. brain_context_pack discloses query_tokens for the text actually sent, instruction prefix included, and its omitted reach now resolves to remote like every other reader.

embedding_input_window_tokens (OPEN_SECOND_BRAIN_EMBEDDING_INPUT_WINDOW_TOKENS, since v1.73.0) declares the input window, in the model's own tokens, of a model the curated table does not list, or overrides the table's value. It is an integer of at least 1; a blank value is refused rather than read as unset. The chunk-window census of an index run and of o2b search status reads the same window, so a declared window enables it for an uncurated model. The offline local embedder has no input window, so the key is refused with INVALID_INPUT when embedding_provider is local.

embedding_extra_body (OPEN_SECOND_BRAIN_EMBEDDING_EXTRA_BODY, since v1.73.0) sends operator-declared fields with every embedding request to an OpenAI-compatible endpoint, for a serving stack that needs a field the provider does not send (a dimensions value, a truncation switch). The value is one JSON object in one key, because the flat config format cannot hold a nested map:

embedding_extra_body: '{"dimensions": 512}'
embedding_dimension: 512

The owned request fields model, input and encoding_format may not be set, in any spelling (case, _ and -, fullwidth and zero-width variants are folded before the comparison), and are refused by name with INVALID_INPUT. A blank value, invalid JSON or a JSON value that is not an object is refused the same way, naming the env variable or the key that supplied it. Only openai-compat sends the body, so the key is refused for any other provider; disabled sends nothing and is exempt. A dimensions field changes the width the provider answers with, so it requires embedding_dimension and must agree with it. The checks run on the resolved config, so a programmatic override meets them too. The extra body is not part of the embedding identity: declaring or editing it never triggers a reindex. Some fields shape the vectors the provider returns (a provider task, input_type, normalize or truncate switch); after changing such a field, rebuild the vectors with o2b search reindex --embeddings so new query vectors stay in the space of the stored passages. Nothing triggers that rebuild on its own, and o2b search index --force is not enough: an unchanged chunk keeps its stored vector.

Vector carry-over (since v1.72.0). When a note is edited, a chunk whose content did not change keeps its stored vector, provided the vector was written by the model and dimension the index records. Only changed chunks are re-embedded and paid for. A paragraph moved within the note is carried too; a paragraph moved into another note is re-embedded. embedding_batch_tokens (OPEN_SECOND_BRAIN_EMBEDDING_BATCH_TOKENS, since v1.43.0) adds a per-request token budget beside embedding_batch_size: a batch closes on whichever cap fills first, so a run of long chunks cannot assemble a request past a provider's per-request token ceiling. The estimate is the same one the cost gate uses, applied to the text as it will be sent (instruction prefix included). With the key unset the field is absent and batching is byte-identical to the fixed embedding_batch_size stride; a single text whose own estimate exceeds the budget is sent alone rather than dropped or split. embedding_concurrency (OPEN_SECOND_BRAIN_EMBEDDING_CONCURRENCY, default 4) bounds embedding requests in flight for one embedding identity - provider, model and configured dimension - against one resolved endpoint, across the whole process rather than one call. Two models configured against the same host are two identities and therefore two budgets, so that host sees up to embedding_concurrency x identities at once: size the knob against the budget the provider actually meters. Two configurations that ask for different values on the same identity and endpoint are refused by name rather than reconciled to one of them. search_fusion_mode (default linear) may be set to rrf to fuse the keyword and semantic lanes by reciprocal rank (search_rrf_k, default 60); linear keeps ranking bit-identical.

FTS tokenizer (since v1.34.0): search_fts_diacritics (default 2; also 0/1) sets the unicode61 remove_diacritics rule, and search_fts_stemmer (default none; also porter) layers Porter stemming over it. Unset keys keep the historical unicode61 remove_diacritics 2 clause byte-identically. An out-of-range value is rejected loudly; the CJK trigram prefilter is unaffected. Changing either key only takes effect after o2b search reindex — there is no implicit reindex.

Typed relations participate in ranking (relation polarity): a page whose frontmatter declares superseded_by: is demoted when it matches and its successor is boosted or pulled in, contradicts: surfaces warning-style why_retrieved reasons on both endpoints without endorsement, and related / extends / depends_on / refines grant a small bounded boost between co-retrieved pages. Vaults without typed relations rank identically; search_relation_polarity_enabled: false (or OPEN_SECOND_BRAIN_SEARCH_RELATION_POLARITY=false) is the kill switch. Learned recall weights are opt-in via search_learned_weights_enabled: true (or OPEN_SECOND_BRAIN_SEARCH_LEARNED_WEIGHTS=true); multipliers stay within [0.8, 1.2] and affected results carry a learned_weights: reason.

Structured recall query documents are line-oriented. intent: accepts neutral, exact, entity, or broad; lex: accepts bare or quoted terms and -excluded tokens; vec: and hyde: provide semantic text lanes when the semantic layer is configured. Example:

intent: entity
lex: "project alpha" -archived
vec: active implementation context
hyde: a note that explains the current project alpha decision

The fused ranking is sharpened by a recall-quality suite (v0.13.0), each layer config-tunable and bounded:

Config key Env var Default Effect
search_mmr_lambda OPEN_SECOND_BRAIN_SEARCH_MMR_LAMBDA 0.7 MMR relevance-vs-diversity tradeoff; 1 disables diversification
search_max_hops OPEN_SECOND_BRAIN_SEARCH_MAX_HOPS 1 Link-graph traversal depth during recall; 0 disables
search_hop_decay OPEN_SECOND_BRAIN_SEARCH_HOP_DECAY 0.5 Per-hop score multiplier for traversal-surfaced docs
search_max_expansion_per_hit OPEN_SECOND_BRAIN_SEARCH_MAX_EXPANSION_PER_HIT 3 Cap on outbound links followed per node

Recall and ranking quality (v0.20.0), each tunable and bounded:

Config key Env var Default Effect
search_recency_shape OPEN_SECOND_BRAIN_SEARCH_RECENCY_SHAPE 0.8 Weibull recency curve shape (k)
search_recency_scale OPEN_SECOND_BRAIN_SEARCH_RECENCY_SCALE 30 Weibull characteristic lifetime in days
search_recency_amplitude OPEN_SECOND_BRAIN_SEARCH_RECENCY_AMPLITUDE 0.05 Max recency boost at age 0; 0 disables the recency layer
search_intent_enabled OPEN_SECOND_BRAIN_SEARCH_INTENT_ENABLED true Re-weight ranking by structural query intent; false is neutral
search_synonym_enabled OPEN_SECOND_BRAIN_SEARCH_SYNONYM_ENABLED false Opt-in co-occurrence query expansion (language-agnostic)
search_synonym_max_terms OPEN_SECOND_BRAIN_SEARCH_SYNONYM_MAX_TERMS 3 Cap on expansion terms OR'd onto the query
search_cache_enabled OPEN_SECOND_BRAIN_SEARCH_CACHE_ENABLED false Opt-in persistent query cache, gated by corpus generation
search_cache_ttl_seconds OPEN_SECOND_BRAIN_SEARCH_CACHE_TTL 300 Cache row time-to-live in seconds
search_chain_stop_enabled OPEN_SECOND_BRAIN_SEARCH_CHAIN_STOP false Opt-in cross-vault early termination once an origin answers confidently
search_chain_stop_score OPEN_SECOND_BRAIN_SEARCH_CHAIN_STOP_SCORE 0.8 Match-quality [0,1] threshold that triggers the chain-stop: the share of the query's IDF mass an origin covered, not its top result score

brain_context_pack also accepts max_chars_per_memory and max_total_chars (code-point caps). Pass --lanes to keep the legacy flat items while also returning directives, constraints, and consider lanes derived from polarity cues and page tier. Surfaced item bodies are guarded by deterministic prompt-injection checks; filtered items return a placeholder and safety.reasons rather than hostile note text. The read-only brain_pre_compress_pack MCP tool returns a budgeted top-preferences-plus-active.md addendum for a host runtime to inject before a context-compression event, with the same safety report shape.

Entity-boosted retrieval and header-anchored chunking populate on the next reindex and need no configuration. Every result carries a why_retrieved list naming the scoring layers that ranked it.

Since v1.64.0 the index persists each document's resolved event-time window (event_time_min / event_time_max, schema v13 - frontmatter validity window, else the event anchor; one NULL side is a declared open window) through the one shared rung order the query side also judges by, so index time and query time cannot drift, and the store answers a window census (declared, intersecting, and the rows still judged by storage mtime) as one SQL aggregate over those bounds. The migration is additive: existing rows keep NULLs - and a NULL window is judged by storage mtime at query time exactly as before - until their next content change or a o2b search reindex refreshes them. A pinned: true page earns a bounded ranking boost (capped at 0.05, reported as a pinned reason and breakdown entry; never an override of the keyword lead). Only the candidates' pinned rows are read, and an unpinned result carries no pinned breakdown key. o2b search status reports the persisted windows on its event_time line.

Two opt-in ranking guards and one deadline join the suite:

Config key Env var Default Effect
search_relational_rerank_pin OPEN_SECOND_BRAIN_SEARCH_RELATIONAL_RERANK_PIN false Rerank may promote a relational-origin candidate but never sink it below its pre-rerank order (the relevance floor is untouched)
search_metadata_boost_gate OPEN_SECOND_BRAIN_SEARCH_METADATA_BOOST_GATE false A query whose keyword lane returned no hits contributes zero from every additive metadata/structural boost layer
search_hybrid_deadline_ms OPEN_SECOND_BRAIN_SEARCH_HYBRID_DEADLINE 15000 Wall-clock budget over the whole composite hybrid path (embed, semantic top-k, rerank, second pass); on expiry the abandoned embed and rerank requests are aborted, the search completes with what it has (keyword-only, or the pre-rerank order with exclusions, trust gate and the other post-rank phases still applied), reports hybridDeadlineExceeded and is not cached; 0 disables
search_freshen_interval_s OPEN_SECOND_BRAIN_SEARCH_FRESHEN_INTERVAL_S 60 Freshen on read (since v1.76.0): a search or a session start that finds the index older than this many seconds starts one background incremental o2b search index at low CPU and I/O priority, and answers from the index it has; 0 turns it off. See "Freshen on read" below
search_freshen_embeddings OPEN_SECOND_BRAIN_SEARCH_FRESHEN_EMBEDDINGS false Let that background run compute embeddings too (it costs money); off, it indexes keyword-only and leaves vectors to an explicit --embeddings run

Both guards are off by default and leave every score byte-identical; the deadline is on by default because its lane budgets already summed to more, and it bounds exactly the phases with no budget of their own.

Freshen on read (since v1.76.0)

Nothing schedules index runs: no daemon, no OS timer, no Hermes job. A search (every reading MCP tool, the recall-inject hook, o2b search query) and a session start (ensureVaultCurrent) read the index's last_indexed_at; when it is older than search_freshen_interval_s (default 60 s, 0 turns it off), the reader answers from the index it has and starts one detached o2b search index --freshen <token>, so the next read sees the current vault. Only agent activity triggers it.

  • One run at a time. The reader takes an exclusive claim, <index dir>/freshen.claim, and skips when another run holds it, an indexer holds the writer lock, or the index directory refuses the claim (a read-only mount, no permission). A claim older than ten minutes belongs to a run that died and is taken over by exactly one reader, under an exclusive takeover lock. A run that still meets another writer ends as a skip, not a failure.
  • Out of the way. The child lowers its CPU priority and runs under ionice -c3 on Linux or taskpolicy -b on macOS when the tool exists. It indexes keyword-only unless search_freshen_embeddings is on.
  • Cheap when nothing changed. A run that added, updated and deleted no document skips link and alias resolution and writes almost nothing. A changed document rewrites only the chunks that changed, so a daily log that grows all day costs its appended tail per run.
  • Failures back off. A failed run records its error in <index dir>/freshen-state.json and the next run waits 1 minute, doubling to an hour. o2b search status shows the last outcome and the backoff; o2b doctor warns with freshen-failing after three failures in a row. Runs that changed something or failed append one row to the index_freshen metrics surface.
  • Staleness is named. A search over an index more than ten minutes old carries the trail code index-stale with detail.ageSeconds. The code names the age only; whether a background run started is not part of it (see the cases above where freshen on read starts none). The code is added when the answer is served, also on a query-cache hit, and is never stored in the cache.
  • Never a foreign vault. Cross-vault and recall-source reads open other indexes read-only and never start a run there.

o2b search watch stays available for a foreground watcher, and the cron recipes of o2b search reindex --cron-template and the maintenance lane are still printed, but keyword freshness no longer needs them; a scheduled run remains the way to compute embeddings periodically.

The reserved visibility token (since v1.54.0)

o2b search query --visibility <token>...   narrow results to pages declaring one of these visibility tokens
o2b search check                           reports how many indexed documents reserve themselves against remote
                                           reads, how many enumerated note-returning surfaces still do not consult
                                           the field, and how many documents hold no frontmatter the index measured

A vault page's visibility: frontmatter carries opaque, language-neutral tokens, and ONE of them is reserved: a page declaring it is not readable at remote reach. Reach is a transport fact rather than a caller argument - stdio and every command in this CLI run at local reach, because the operator is running this binary in their own shell against a vault they can already open in an editor, while an HTTP bind is local only on loopback. A CLI that withheld the operator's own reserved pages from the operator would hide data from the only party entitled to it while proving nothing to anyone.

The rule is enforced at three read roots - the ranked search pipeline, the page walker, and the by-path and by-chunk-id read primitives - and an architecture sweep fails the suite when a file under src/mcp/, src/cli/ or src/openclaw/ reads a vault path directly without a registered reason. --visibility keeps its existing meaning and can only NARROW: it selects among the other tokens and cannot lift the reserved one. The reserved token also dominates the other tokens on the same page: a page tagged visibility: [private, team] is not returned for --visibility team (before this fix, any one matching token lifted it). A vault that never wrote the reserved token behaves exactly as it did before.

What this boundary is, precisely: it is enforced against callers at remote reach (a non-loopback HTTP bind). An agent on stdio or the CLI runs at local reach, which is filesystem-equivalent access to the vault, so the reserved token does not bind it beyond the --visibility rule above. Whole-vault generated views (osb://preferences/active, osb://lessons, osb://digest/latest, osb://status, and the brief and digest views that render them) are not filtered by reach; o2b search check counts them among the surfaces that do not consult the field.

Fail-closed, in both directions that matter. A page whose file cannot be read - deleted, renamed, moved to a failing mount, or chmod'd between one call and the next - is treated as declaring the reserved token at remote reach, because an unreadable visibility claim is not the absence of one; at local reach it is kept, since the operator is the party who has to fix it. A withheld page is reported exactly as an absent one, and no result count on a read surface states how many rows were withheld - the only place the withheld population is named is the o2b search check diagnostic above, which the operator ran deliberately at local reach.

Schema 12 adds documents.visibility, holding what the INDEX measured of each page's frontmatter, and backfills it during the migration from the frontmatter the index already stores - one table scan, no vault walk, no reindex. The column reports; the live file decides. Three states are kept distinct: the page's tokens, an empty list for a page that declares none, and unmeasured for a document the index holds no frontmatter chunk for. o2b search check reports the unmeasured population rather than counting it as declaring nothing.

What leaves the machine unscanned (network egress)

Every export that writes a FILE - brain export, brain explorer --export, brain bank-export, brain graph-export, brain okf-export, export-config, install --out - runs the shared redactor over its bytes before they land, and prints a notice on stderr when it removed something.

The widest of them is brain export --format transcripts-jsonl (since v1.50.0), and it is worth naming apart from the other six. The rest carry vault artifacts an agent composed; this one carries whole recorded conversations from whichever runtimes wrote them, which is where a key pasted into a prompt actually lives. Three properties follow from that:

  • its records are guarded one at a time, so a single oversized turn cannot push a machine's whole corpus past the redactor's scan window;
  • it is the one call site that judges identifiers as FOREIGN. session_id and turn_id were named by the harness that wrote the transcript, so the guard's default narrowing - an argument about ids this vault constructs - does not cover them, and the full bare-token detector is used instead;
  • a secret-shaped identifier refuses the whole export and writes nothing the operator can see, because the only file that exists at that point is an unnamed spool and it is deleted.

The registry entry is brain-export (renamed from brain-preference-export when the third format joined), and it covers all three of the verb's formats: the JSON and transcript forms are redacted as TREES and serialised afterwards, llms-txt as text.

Five paths send vault-derived bytes to a NETWORK destination and are not scanned. That is a stated exposure, not an oversight: redacting an embedding input produces a vector for the placeholder rather than for the text, so the chunk comes back unfindable while the index reports success. The full decision record for each is src/core/egress/registry.ts, keyed by the id below.

id verb what leaves when it can happen
search-embedding-openai-compat o2b search index / any reindex every indexed chunk BODY, verbatim, plus any fields declared in embedding_extra_body only once search_embedding_endpoint + an API key are configured; the endpoint is whichever host you name, including a local one
search-embedding-zeroentropy o2b search index (zeroentropy profile) the same chunk bodies, to a second vendor's embed endpoint same gate, when that provider profile is selected
search-rerank-cross-encoder o2b search --rerank the QUERY plus the top-of-pool candidate DOCUMENTS - vault text the embedding path may never have seen, chosen by relevance to what you just asked only with a reranker endpoint configured
brain-telegram-capture o2b brain telegram-run reply text POSTed to the Telegram Bot API; the /catchup reply is composed from vault content only while the runner verb is running; an install that never starts it never reaches this path
research-external-fetch o2b brain research an agent-composed search query (not a page body) to the configured research provider key-gated: with no key set every call is a typed disabled error

Semantic search, reranking, Telegram capture and research are all off until you configure an endpoint, so a default install has no network egress at all.

Endpoint scheme. The embedding (embedding_base_url) and reranker (search_rerank_base_url) endpoints must be https://; plain http:// is accepted only for a loopback host (localhost, 127.0.0.1, ::1), and none of the provider requests follows a redirect. A local server on another machine - LM Studio on the Windows host seen from WSL, Ollama on a LAN box, anything at a tailnet 100.x address - has no certificate to offer, so each endpoint has its own explicit opt-out:

embedding_base_url: http://100.64.0.5:1234/v1
embedding_allow_insecure_http: true        # OPEN_SECOND_BRAIN_EMBEDDING_ALLOW_INSECURE_HTTP
search_rerank_base_url: http://192.168.1.20:8080/v1
search_rerank_allow_insecure_http: true    # OPEN_SECOND_BRAIN_SEARCH_RERANK_ALLOW_INSECURE_HTTP

Both default to false. An opt-out applies only to a base URL set in the o2b config or environment, never to one supplied by a provider profile registered with o2b search provider add (that registry lives inside the vault, and a vault write must not be able to point an opt-out at a new host). A process that uses one prints a warning on stderr once per endpoint: chunk text, the query and the API key travel unencrypted, so use it only on a network you trust. tests/core/architecture/egress-census.test.ts fails if a sixth such path is added without a declaration, and fails if a declared one is missing from this table.

One network path IS scanned: decision-model-systemone, the optional decision model (docs/decision-models.md). It is off unless the operator enables it in machine config and the named key variable is set. What leaves is a masked, clipped state (for a rerank, the query plus the top candidate passages as P0..Pn, never a page whose visibility is private or cannot be resolved, nor a chunk that carries part of a <private> region) and the question texts; the whole body passes redactForEgress, and a refused body is not sent. The same holds for decision-model-llm-emulation, the optional uncalibrated route that sends that state to an OpenAI-compatible chat model and only when the operator names it (docs/decision-models/providers.md).

Decision model (optional)

o2b decision-model check      Config state: enabled, provider, base URL, pinned model, the key
                              variable's NAME and whether it is set (never the value), per-use
                              modes, vault opt-out, cost gate and today's spend, processor terms,
                              adapter and calibration, threshold profile (with one warning for
                              uses whose enforce runs as shadow), choice option limit, licence
                              note for open weights with a commercial-use restriction.
                              --ping sends one request over a synthetic state (no vault content)
                              and records it (use ping, counted toward the cost gate).
                              Exit 1 only for an invalid config or a failed ping of an active
                              provider; a missing key is exit 0 with a hint. --vault --config --json
o2b decision-model report     decision_model_call records per use: calls, outcome mix, p50/p95
                              latency, input tokens, cost, rerank shadow agreement (top-1, top-5
                              overlap, ordinary shadow records only). --since DATE --use USE
                              --vault --config --json

Off by default; active only when decision_model_enabled: "true" AND the environment variable named by decision_model_env_key is set. Config keys, presets, privacy and the vault opt-out: docs/decision-models.md.

Helpers

o2b-hook                      Internal launcher invoked by hooks/hooks.json (Claude Code & Codex)
vault-log                     Shell mirror of brain_note (one-liner narrative milestones)

Conventions

  • Every CLI mutation that touches Brain takes a pre-run snapshot under Brain/.snapshots/ with a SHA-256 sidecar manifest. o2b brain rollback aborts on drift unless --force-rollback.
  • o2b ... --json exists on every read verb and most write verbs (the JSON payload mirrors the MCP tool's response shape).
  • Root-level --json is inherited by every command parser; existing semantic JSON commands keep their native payloads, while other commands emit a redacted fallback envelope.
  • MCP-only brain_pinned_context manages Brain/pinned.md, a transient current-task scratchpad loaded by brain_context; it is intentionally not a learned preference CLI verb.
  • --dry-run is supported by every mutating verb that touches more than a single file (brain merge, brain rollback, brain upgrade, brain import-claude-memory, update, ...).
  • --vault always overrides the profile path; useful for multi-vault hosts where the config-resolved default is not the right target.