Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions .agents/skills/harness-adapters/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -120,19 +120,21 @@ Use `low` for well-understood work with an explicit bounded path and `xhigh` for
Choose intermediate levels proportionally as complexity, uncertainty, blast radius, or open-ended reasoning increases.
When a verified adapter lacks `xhigh`, cap the choice at its highest supported non-`max` level rather than omitting the intended effort silently.
Never select `max` from this fallback; use it only when the captain has explicitly expressed that per-task or standing preference.
Never select Codex-only `ultra` from this fallback.
Before selecting `ultra`, explain that it uses internal sub-agent decomposition with significantly higher token spend per turn and obtain the captain's current explicit approval.

The supported launch-profile flags below are verified locally; each row records its evidence.

| Harness | Model flag | Effort flag | Notes |
|---|---|---|---|
| claude | `--model <model>` | `--effort <low\|medium\|high\|xhigh\|max>` | Verified on Claude Code 2.1.196. |
| codex | `--model <model>` | `-c 'model_reasoning_effort="<low\|medium\|high\|xhigh>"'` | Verified on codex-cli 0.142.1. The installed binary schema contains `model_reasoning_effort`, the active config uses it, and the bundled model catalog advertises only low/medium/high/xhigh. `max` is omitted. |
| codex | `--model <model>` | `-c 'model_reasoning_effort="<low\|medium\|high\|xhigh\|ultra>"'` | Verified on codex-cli 0.142.1 for low/medium/high/xhigh: the installed binary schema contained `model_reasoning_effort`, the active config used it, and the bundled model catalog advertised those four. codex-cli 0.149.1's embedded schema adds `ultra`, exclusive to gpt-5.6-sol. The PATH codex is currently nix-pinned at 0.133.0 and does not support `ultra` until its separate cutover, so an `ultra` launch fails there. `max` is omitted. |
| grok | `--model <model>` | `--reasoning-effort <low\|medium\|high>` | Verified on grok 0.2.99 (2026-07-13). `--effort` is an alias, but firstmate's profile axis is reasoning effort. As of 0.2.99 the ceiling is `high`; both `xhigh` and `max` are rejected with `use one of: high, medium, low`, so firstmate omits them. |
| pi / pi-signed | `--model <model>` | `--thinking <low\|medium\|high\|xhigh\|max>` | Verified 2026-07-27 on Pi and pi-signed 0.82.0. Both expose the same accepted thinking levels and completed the same model-qualified max-thinking smoke. |
| opencode | `--model <provider/model>` | none for firstmate's interactive launch | Verified on opencode 1.17.6. `opencode run` has `--variant`, but firstmate launches the interactive `opencode --prompt` path, which has no verified effort flag. |
| kimi | `--model <model>` | none | Verified 2026-07-25 on Kimi Code CLI 0.29.1. |
| cursor | `--model <model>` | none | Verified 2026-08-11 on Cursor Agent CLI 2026.08.11-e8db854. No effort flag exists, so firstmate records the requested effort in task metadata and omits it from the launch. Validate ids against `cursor-agent --list-models` rather than assuming a low/medium/high family: the live catalog carries only `-high` Grok ids. |
| muse | `--model <model>` | `--reasoning-effort <low\|medium\|high\|xhigh>`, and `ultra` only for an explicit `max` | Verified 2026-08-05 on Muse Code 0.1.0-R708.1. The flag accepts `none\|minimal\|low\|medium\|high\|xhigh\|ultra` and defaults to `high`. `ultra` is muse's max-class level, so it is reachable only through an explicit captain `max`, never from the generic fallback; `none` and `minimal` sit below the shared vocabulary and stay unreachable. |
| muse | `--model <model>` | `--reasoning-effort <low\|medium\|high\|xhigh>`, and muse-native `ultra` only for an explicit `max` | Verified 2026-08-05 on Muse Code 0.1.0-R708.1. The flag accepts `none\|minimal\|low\|medium\|high\|xhigh\|ultra` and defaults to `high`. Muse's native `ultra` is its own max-class level and is unrelated to firstmate's Codex-only `ultra` profile value: firstmate reaches muse's level only by mapping an explicit captain `max` onto it, never from the generic fallback, while a profile carrying the Codex-only `ultra` is recorded in task meta and omitted from the muse launch, leaving muse on its `high` default. `none` and `minimal` sit below the shared vocabulary and stay unreachable. |

The concrete `harness` field owns adapter identity independently of the model provider: `harness=pi` with `model=xai/grok-*` is Pi using xAI, not `harness=grok`, and does not require Grok CLI login; `harness=grok` remains the standalone Grok Build CLI adapter.
Likewise, `harness=cursor` with `model=cursor-grok-4.5-*` is Cursor Agent CLI routing a Grok model, not the xAI Grok Build `grok` harness.
Expand Down
2 changes: 1 addition & 1 deletion bin/fm-bootstrap.sh
Original file line number Diff line number Diff line change
Expand Up @@ -1085,7 +1085,7 @@ crew_dispatch_validate() {
if $e == null then true
elif ($e | type) != "string" then false
elif $h == "claude" then (["low","medium","high","xhigh","max"] | index($e))
elif $h == "codex" then (["low","medium","high","xhigh"] | index($e))
elif $h == "codex" then (["low","medium","high","xhigh","ultra"] | index($e))
elif $h == "grok" then (["low","medium","high"] | index($e))
elif $h == "pi" or $h == "pi-signed" then (["low","medium","high","xhigh","max"] | index($e))
elif $h == "muse" then (["low","medium","high","xhigh","max"] | index($e))
Expand Down
8 changes: 4 additions & 4 deletions bin/fm-control.sh
Original file line number Diff line number Diff line change
Expand Up @@ -243,8 +243,8 @@ fi
[ "$MODEL_SET" = 0 ] || [ -n "$NEW_MODEL" ] || die "--model requires a non-empty value"
[ "$EFFORT_SET" = 0 ] || [ -n "$NEW_EFFORT" ] || die "--effort requires a non-empty value"
case "$NEW_EFFORT" in
''|low|medium|high|xhigh|max) ;;
*) die "--effort must be one of low, medium, high, xhigh, max" ;;
''|low|medium|high|xhigh|max|ultra) ;;
*) die "--effort must be one of low, medium, high, xhigh, max, ultra" ;;
esac

# --- exact task-id resolution ----------------------------------------------
Expand Down Expand Up @@ -632,9 +632,9 @@ resolve_relaunch_profile() {
CONFIG_MODEL=$("$SCRIPT_DIR/fm-harness.sh" secondmate-model 2>/dev/null || true)
CONFIG_EFFORT=$("$SCRIPT_DIR/fm-harness.sh" secondmate-effort 2>/dev/null || true)
case "$CONFIG_EFFORT" in
''|low|medium|high|xhigh|max) ;;
''|low|medium|high|xhigh|max|ultra) ;;
*)
echo "warning: config/secondmate-harness effort token '$CONFIG_EFFORT' is not one of low, medium, high, xhigh, max; ignoring" >&2
echo "warning: config/secondmate-harness effort token '$CONFIG_EFFORT' is not one of low, medium, high, xhigh, max, ultra; ignoring" >&2
CONFIG_EFFORT=
;;
esac
Expand Down
2 changes: 1 addition & 1 deletion bin/fm-remote-secondmate-control.sh
Original file line number Diff line number Diff line change
Expand Up @@ -144,7 +144,7 @@ cmd_launch() {
claude|codex|opencode|pi|pi-signed|grok|kimi|cursor) ;;
*) die "unverified remote secondmate harness: $harness" ;;
esac
case "$effort" in -|low|medium|high|xhigh|max) ;; *) die "invalid remote secondmate effort: $effort" ;; esac
case "$effort" in -|low|medium|high|xhigh|max|ultra) ;; *) die "invalid remote secondmate effort: $effort" ;; esac
# Herdr is required on this host, not merely preferred: its server belongs to
# the GUI login session, so the endpoint survives every SSH disconnection that
# a remote route depends on. bin/fm-remote-doctor.sh is the readiness owner.
Expand Down
45 changes: 26 additions & 19 deletions bin/fm-spawn.sh
Original file line number Diff line number Diff line change
Expand Up @@ -34,10 +34,11 @@
# the new incarnation.
# --harness <name> is the explicit per-spawn harness/profile adapter. The old
# positional harness arg still works for back-compat.
# --model <name> and --effort <low|medium|high|xhigh|max> are concrete profile
# axes chosen by firstmate at intake. They are only threaded into harnesses whose
# installed CLIs were verified to support that axis; unsupported axes are omitted
# from that harness's launch rather than guessed.
# --model <name> and --effort <low|medium|high|xhigh|max|ultra> are concrete
# profile axes chosen by firstmate at intake. ultra is Codex-only. Values are
# threaded only into harnesses whose installed CLIs were verified to support
# them; unsupported axes are omitted from that harness's launch rather than
# guessed.
# --backend <name> is the explicit runtime session-provider backend for this
# exact task only (docs/configuration.md "Runtime backend" owns when that flag
# is authorized). Without it, the script resolves FM_BACKEND, then
Expand Down Expand Up @@ -346,8 +347,8 @@ if [ "$TRACEPARENT_SET" -eq 1 ]; then
}
fi
case "$EFFORT" in
''|low|medium|high|xhigh|max) ;;
*) echo "error: --effort must be one of low, medium, high, xhigh, max" >&2; exit 1 ;;
''|low|medium|high|xhigh|max|ultra) ;;
*) echo "error: --effort must be one of low, medium, high, xhigh, max, ultra" >&2; exit 1 ;;
esac

# --relaunch reuses an existing task's endpoint, worktree, project, and kind,
Expand Down Expand Up @@ -473,7 +474,7 @@ spawn_remote_secondmate() {
;;
esac
case "$effort" in
-|low|medium|high|xhigh|max) ;;
-|low|medium|high|xhigh|max|ultra) ;;
*)
fm_lock_release "$registry_lock" || true
fm_lock_release "$SPAWN_TASK_LOCK" || true
Expand Down Expand Up @@ -1288,8 +1289,8 @@ if [ "$KIND" = secondmate ] && [ -z "$ARG3" ]; then
SM_EFFORT=$("$SCRIPT_DIR/fm-harness.sh" secondmate-effort)
if [ -n "$SM_EFFORT" ]; then
case "$SM_EFFORT" in
low|medium|high|xhigh|max) EFFORT=$SM_EFFORT ;;
*) echo "warning: config/secondmate-harness effort token '$SM_EFFORT' is not one of low, medium, high, xhigh, max; ignoring" >&2 ;;
low|medium|high|xhigh|max|ultra) EFFORT=$SM_EFFORT ;;
*) echo "warning: config/secondmate-harness effort token '$SM_EFFORT' is not one of low, medium, high, xhigh, max, ultra; ignoring" >&2 ;;
esac
fi
fi
Expand Down Expand Up @@ -1392,11 +1393,13 @@ effort_flag_for_harness() {
esac
;;
codex)
# The installed codex config schema uses model_reasoning_effort, and the
# bundled model catalog advertises low|medium|high|xhigh. Omit max rather
# than passing an unsupported value.
# codex-cli 0.149.1's schema accepts model_reasoning_effort values
# low|medium|high|xhigh|ultra. ultra is exclusive to gpt-5.6-sol, and the
# currently nix-pinned PATH codex 0.133.0 rejects it outright until its
# separately owned cutover. Omit max rather than passing an unsupported
# value.
case "$effort" in
low|medium|high|xhigh) printf -- '-c %s ' "$(shell_quote "model_reasoning_effort=\"$effort\"")" ;;
low|medium|high|xhigh|ultra) printf -- '-c %s ' "$(shell_quote "model_reasoning_effort=\"$effort\"")" ;;
esac
;;
grok)
Expand All @@ -1409,19 +1412,23 @@ effort_flag_for_harness() {
esac
;;
pi|pi-signed)
# Pi 0.80.6 accepts the full shared effort vocabulary, including max, through
# its --thinking flag.
# Pi 0.80.6 accepts low through max through its --thinking flag. Pi has no
# ultra concept and its ladder ends at max, so the Codex-only ultra is
# omitted rather than passed to a flag that rejects it.
case "$effort" in
low|medium|high|xhigh|max) printf -- '--thinking %s ' "$(shell_quote "$effort")" ;;
esac
;;
muse)
# muse 0.1.0-R708.1 --reasoning-effort accepts none|minimal|low|medium|
# high|xhigh|ultra and defaults to high, so low..xhigh map straight across.
# ultra is muse's max-CLASS level, so firstmate's max maps onto it - but
# only ever as an EXPLICIT captain choice, never as a fallback, because
# AGENTS.md section 4 forbids selecting max without captain preference and
# the omitted effort here leaves muse on its own high default. muse's extra
# muse's native ultra is muse's own max-CLASS level and is unrelated to the
# Codex-only ultra profile value emitted above. firstmate's max maps onto
# muse's level, but only ever as an EXPLICIT captain choice, never as a
# fallback, because AGENTS.md section 4 forbids selecting max without
# captain preference and the omitted effort here leaves muse on its own
# high default. The Codex-only ultra is therefore deliberately absent from
# the case below and falls through to that same high default. muse's extra
# none/minimal levels sit below firstmate's shared vocabulary and are
# deliberately unreachable rather than remapped onto low.
case "$effort" in
Expand Down
1 change: 1 addition & 0 deletions bin/fm-test-run.sh
Original file line number Diff line number Diff line change
Expand Up @@ -173,6 +173,7 @@ family_for_basename() {
fm-backlog-handoff.test.sh|fm-on.test.sh|fm-remote-backlog-handoff.test.sh|\
fm-remote-doctor.test.sh|fm-remote-job.test.sh|fm-remote-job-orphan-reap.test.sh|\
fm-remote-reply.test.sh|fm-remote-secondmate-lifecycle-e2e.test.sh|\
fm-remote-secondmate-profile-axes.test.sh|\
fm-remote-secondmate-trace-context.test.sh|\
fm-secondmate-harness.test.sh|fm-secondmate-lifecycle-e2e.test.sh|\
fm-secondmate-liveness.test.sh|fm-secondmate-safety.test.sh|fm-secondmate-sync.test.sh|\
Expand Down
6 changes: 5 additions & 1 deletion docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -288,6 +288,7 @@ When it is absent or contains `default`, crewmates mirror the firstmate's own ha
`config/secondmate-harness` is a separate local, gitignored file containing the adapter the primary uses to launch secondmate agents, optionally followed by model and effort tokens on the same line.
The first non-empty, non-comment line is parsed as `<harness> [<model>] [<effort>]`.
A bare `<harness>` preserves the previous behavior: harness only, with no model or effort launch flag.
The effort token accepts the Codex-only `ultra` as well, and [Crew dispatch profiles](#crew-dispatch-profiles-configcrew-dispatchjson) below owns its Codex pairing and codex-cli floor.
When the harness token is absent or `default`, secondmate launch falls back through `config/crew-harness` and then the primary's own harness, and no model or effort is read from that file.
`fm-harness.sh secondmate-model` and `fm-harness.sh secondmate-effort` expose only the optional tokens from `config/secondmate-harness`; `config/crew-harness` remains a bare adapter-name file.
Changing this pin affects the next secondmate spawn or control-plane relaunch; the relaunch profile rules are owned by [`docs/agent-control.md`](agent-control.md#transactional-relaunch).
Expand Down Expand Up @@ -321,7 +322,7 @@ This section is the single owner of the canonical schema and its per-field seman
{
"when": "<natural-language condition describing a kind of task>",
"use": [
{ "harness": "<adapter>", "model": "<optional model>", "effort": "<low|medium|high|xhigh|max, optional>" }
{ "harness": "<adapter>", "model": "<optional model>", "effort": "<low|medium|high|xhigh|max, or codex-only ultra; optional>" }
],
"why": "<optional rationale that helps firstmate choose>"
}
Expand All @@ -337,6 +338,9 @@ Both `use` and the optional top-level `default` accept either one profile object
The single-object form stays fully backward-compatible, and every profile needs `harness`.
Profile `model` and `effort` fields and rule `why` are optional.
An omitted model or effort means the selected harness uses its own default for that axis.
`ultra` is a Codex-only profile value, and bootstrap validates that harness pairing while every other harness keeps its existing effort set.
Codex supports `ultra` only on `gpt-5.6-sol`, so select that model when using it.
`ultra` also needs codex-cli 0.149.1 or newer on `PATH`; the PATH codex is currently nix-pinned at 0.133.0, which does not support it, so an `ultra` profile fails at pane launch until that CLI is rebuilt and cut over.
Every profile array is an implicit quota-aware choice resolved through `quota-array-dispatch`.
If no dispatch rule fits, firstmate resolves `default` through the same object-or-array path before falling back to `config/crew-harness`.
If a selected profile carries an effort value the chosen harness does not accept, `fm-spawn.sh` records the requested `effort=` in task meta for traceability but omits the launch flag, and bootstrap reports the invalid harness/effort pair as a `CREW_DISPATCH` diagnostic when it is visible in the file.
Expand Down
3 changes: 3 additions & 0 deletions docs/remote-secondmates.md
Original file line number Diff line number Diff line change
Expand Up @@ -251,9 +251,12 @@ bin/fm-test-run.sh tests/fm-project-origin.test.sh
bin/fm-test-run.sh tests/fm-remote-reply.test.sh
bin/fm-test-run.sh tests/fm-remote-backlog-handoff.test.sh
bin/fm-test-run.sh tests/fm-remote-secondmate-lifecycle-e2e.test.sh
bin/fm-test-run.sh tests/fm-remote-secondmate-profile-axes.test.sh
bin/fm-test-run.sh tests/fm-remote-secondmate-trace-context.test.sh
```

`tests/fm-remote-secondmate-profile-axes.test.sh` owns the launch-profile axes on this route: the harness, model, and effort a `config/secondmate-harness` line pins are validated by the parent, replayed across the SSH boundary by `bin/fm-remote-secondmate-control.sh`, and applied by the remote host's own `bin/fm-spawn.sh`, so the launch literal the remote pane receives is what the suite reads back.

The account-level checks the doctor performs - a real Aqua login session, a real `launchctl` domain, and a real herdr server - are only ever exercised against fixtures here, so the readiness gate's behavior on a genuine Mac remains an operator-run smoke test.

For a real-host smoke test, provision a disposable remote account and project, run the doctor and its repair against that account, launch the second mate, send one marked request, verify its correlated reply and structured fleet projection, simulate an unreachable host to confirm unknown-without-failover behavior, then retire only after the remote queue is empty.
Expand Down
2 changes: 2 additions & 0 deletions tests/fm-bootstrap.test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -1119,6 +1119,8 @@ test_crew_dispatch_validation() {
malformed dispatch config is flagged^{"rules":[^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - malformed JSON
unverified dispatch harness is flagged^{"rules":[{"when":"anything","use":{"harness":"spaceship"}}],"default":{"harness":"codex"}}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - unverified harness: spaceship
unsupported codex max effort is flagged^{"rules":[{"when":"big feature","use":{"harness":"codex","model":"gpt-5","effort":"max"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: codex:max
codex ultra effort is accepted^{"rules":[{"when":"ultra coding","use":{"harness":"codex","model":"gpt-5.6-sol","effort":"ultra"}}]}^empty^
unsupported claude ultra effort is flagged^{"rules":[{"when":"ultra coding","use":{"harness":"claude","model":"claude-opus-4-6","effort":"ultra"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: claude:ultra
unsupported grok max effort is flagged^{"rules":[{"when":"deep current work","use":{"harness":"grok","model":"grok-4","effort":"max"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: grok:max
unsupported grok xhigh effort is flagged^{"rules":[{"when":"deep current work","use":{"harness":"grok","model":"grok-4","effort":"xhigh"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: grok:xhigh
pi max effort is accepted^{"rules":[{"when":"deep coding","use":{"harness":"pi","model":"openai-codex/gpt-5.6-sol","effort":"max"}}]}^empty^
Expand Down
Loading