Skip to content

feat: make the default backend configurable with session.defaultProvider - #344

Merged
saucam merged 6 commits into
mainfrom
feat/default-provider-339
Sep 27, 2026
Merged

saucam merged 6 commits into
mainfrom
feat/default-provider-339

Conversation

@saucam

@saucam saucam commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #339

What this changes

codeoid always put a session on Claude unless the caller named another backend.
An operator who runs only pi, codex or qwen had to pass --provider on every session.

The default is now a setting: session.defaultProvider in config.json, or CODEOID_DEFAULT_PROVIDER.
It sits next to session.defaultModel, which is where the original multi-provider design put it.
The issue proposed a top-level key instead.

{ "session": { "defaultProvider": "pi" } }

It fails loudly instead of quietly

  • At startup: a misspelled, disabled or uninstalled default stops the daemon with one line naming the fix, e.g. set providers.geminiCli.enabled to true, or the missing API key.
    It does not fall back to Claude.

  • In Settings: a save is refused if the next boot would fail to start.
    The daemon checks the config that boot would load, so it catches:

    • a bad default;
    • disabling the default backend;
    • clearing the default backend's API key;
    • a value that only fails together with an env override.

    If the config is already broken, for example because the default's binary disappeared after an upgrade, unrelated saves still go through, so Settings can still be used to repair it.
    A refusal is shown under the field that caused it.

What else had to change for the default to be safe

  • Model checks: four code paths treated "no provider given" as Claude even though the session would run on the default.
    With a pi default, a new pi session was checked against Claude's models, so model: opus was accepted.
    Those paths now use the actual default.
  • Old sessions: a session saved before multi-backend support, or one whose backend has since gone, resumes on Claude, not on the new default.
    Moving it to another backend would lose its conversation.
  • Queued workers: a worker task queued without a backend is pinned to the backend its model was checked against.
  • Telegram /model: it lists the attached session's models, not the default backend's.
  • Unchanged: the conductor stays on conductor.provider, because it needs the fleet tools, which only Claude provides today.
    Startup logs which backend is the default.
    The settings help notes that backends gate tool use differently.

Verification

  • Tests: 37 new (one obsolete test removed), covering the registry, config, settings, resume, dispatch, Telegram and collaboration resume.
    Each one fails when its fix is reverted, checked by temporarily reverting each fix.

  • Suites:

    • daemon: bun test 2656 pass, 0 fail;
    • web: 571 pass;
    • typecheck and lint clean.
  • Live run of a real daemon from this branch with an isolated config:

    • A misspelled default in config.json: the daemon exited 1 with the one-line error.
    • CODEOID_DEFAULT_PROVIDER=codex:
      • codex was listed first;
      • models.list served codex;
      • a provider-less session landed on codex and answered a real turn;
      • model: opus was refused for codex;
      • a pre-upgrade session resumed on Claude;
      • Settings refused claud and accepted codex.
  • Audit: four rounds, each a correctness review plus a security review.

    • Round 1: ship with fixes, security GO.
    • Round 2: ship with fixes, security GO.
    • Round 3: ship, security GO.
    • Round 4: ship with one fix, security GO.
    • Final audit of the whole diff: ship; its one stale doc line is fixed.

    Every finding was fixed in this branch, except the one below.

  • Not run: the Highflame regression suite, because codeoid isn't part of it.
    The live run above stands in for it.

  • Known gap, left as a follow-up:

    • The trigger is an env override (CODEOID_DEFAULT_PROVIDER) that masks the operator's own config.json default, while a save disables the backend config.json names.
    • That save goes through, and a boot without the override would then fail.
    • Closing it would double the check's complexity for a setup the operator created on purpose.
For agents: where the code is
  • src/config.ts:
    • SessionSchema.defaultProvider (trimmed, non-empty) and DEFAULT_PROVIDER_ENV.
    • loadConfig({ raw, quiet }) loads an in-memory config for previews.
  • src/daemon/providers/registry.ts:
    • createDefaultProviderRegistry(config, env = process.env) takes the default from config and its env as a parameter, for dry runs.
    • It throws DefaultProviderError after registration.
    • defaultProviderProblem() has an id-to-config-key map for "disabled" hints.
  • src/cli.ts: catches only DefaultProviderError, prints one line and exits 1.
  • src/daemon/session-manager.ts:
    • DEFAULT_PROVIDER_ID (always Claude) is removed. "No provider" sites (models.list, create model validation, pack-role validation, dispatch resolveBackend) use #providers.defaultId. Claude-specific sites (the fallback catalog, the conductor) use CLAUDE_PROVIDER_ID.
    • #resumeProviderId: a missing or unregistered persisted provider becomes claude. It is used for both the Session and #resumeRoleChild, and the planned model is dropped when the backend differs.
    • #nextBootProblem / #bootProblem / #culpritKey: the settings next-boot check. It runs on every batch using previewPatches + loadConfig({raw, env}) + a registry build.
      • A batch that sets the default is always checked without the env override.
      • Everything else follows the before/after rule.
      • A refusal is audited as reason=next-boot.
  • src/daemon/settings/store.ts: previewPatches() merges patches onto config.json and a copy of process.env without writing.
  • src/frontends/telegram/index.ts: /model passes the attached session's provider.
  • Tests:
    • src/tests/default-provider.test.ts (new): manager level, with a real registry and mock factories.
    • Also provider-registry, config, models, telegram-flows and collaboration.

🤖 Generated with Claude Code

saucam and others added 5 commits September 26, 2026 19:09
The registry default was the literal "claude", so an operator running only
pi, codex or qwen had to pass --provider on every session. It now comes
from session.defaultProvider (env CODEOID_DEFAULT_PROVIDER).

- Fails fast: a misspelled, disabled or uninstalled default stops daemon
  startup with one actionable line, instead of resolve()'s warn-and-fall-back
  (which exists for resuming sessions from a newer codeoid).
- settings.set checks the value against the live registry with the same
  rule, so the settings UI can't save a default the next boot refuses.
- "No provider" now means the registry default everywhere it meant
  claude: models.list, session.create model validation, pack-role model
  validation, and dispatch model canonicalisation. DEFAULT_PROVIDER_ID,
  which always meant claude, is gone.
- A resumed meta with no providerId predates multi-backend support and stays
  on claude, rather than moving to the new default and losing its backing
  conversation. The conductor keeps conductor.provider.
- Telegram /model lists the attached session's catalog instead of the
  default backend's.

Closes #339

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Audit round 1 on #339.

- settings.set validated session.defaultProvider against the LIVE
  registry, so "default pi + disable pi" in one batch, or disabling or
  clearing the key of an existing default later, was saved and the next
  boot refused to start. It now dry-runs createDefaultProviderRegistry
  on the post-batch config and env (loadConfig takes an in-memory object;
  the registry takes its env as a parameter) for any batch touching the
  default, providers.*, or an env key. Enabling a backend and making it
  the default in one batch is accepted, where the live check refused it.
  Refused writes are audited.
- A resumed session whose backend is no longer available resumes on
  claude, not on the configured default.
- A provider-less fleet spawn whose model was checked against the default
  is pinned to that backend, so a later default change can't put its model
  on another vendor.
- The "disabled" hint names the real switch (providers.geminiCli.enabled)
  and is only offered for backends that have one; the id is JSON-quoted.
- Startup logs which backend is the default; the settings help notes that
  backends' approval gates differ.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ry one

Audit round 2 on #339.

- The next-boot check refused ANY save touching an env key or providers.*
  once the default was already broken (its binary gone after an upgrade),
  and blamed the unrelated key, turning Settings into a wall exactly when it
  was needed. It now refuses only what the batch causes: a batch that sets
  the default is always checked, any other only if the config boots now and
  wouldn't after it.
- An env override that makes the config fail to load is refused the same
  way; schema failures are still applyPatches' to report, per field.
- The check runs inside the handler's error wrapper, and its preview load is
  quiet (no startup warnings repeated on every save).
- Resume resolves a session's backend once and uses it for the role rebuild
  too: a planned collaboration child whose backend is gone resumes on claude
  without carrying the other vendor's model.
- Settings tests no longer depend on the shell's CODEOID_DEFAULT_PROVIDER or
  the bundled pi resolving.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lt on its own

Audit round 3 on #339.

- A config value can be valid alone and fail against an env override
  (mcpToolTimeoutMs saved against a CODEOID_TURN_STALL_TIMEOUT_MS in .env),
  which the next boot rejects. Every batch now gets the next-boot load
  check, not only provider-shaped ones; the before/after rule still lets
  unrelated saves through while the config is already broken, and an
  error while checking the current config counts as "already broken".
- A default the batch sets is also checked with CODEOID_DEFAULT_PROVIDER
  taken out, so an env override can't mask a typo that config.json would
  carry into a later boot.
- "Always check a batch that sets the default" now covers backend
  problems only; a pre-existing load problem follows the before/after rule
  and is no longer blamed on the default.
- Test hygiene: env keys a test saves are restored.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Audit round 4 on #339.

- A refused multi-key batch was filed under its first key. The web drawer
  sends every tab's edits in one batch and shows an error only beside the
  matching field, so the error could land on another tab and the save
  looked like a no-op. The culprit is now the patch whose removal makes the
  next boot work, else "" (the drawer's save bar).
- The full-env check on a batch that sets the default was redundant, and
  its only remaining effect was blaming the default for a broken
  CODEOID_DEFAULT_PROVIDER the batch didn't cause. Dropped; that case follows
  the before/after rule.
- Load-error messages returned to the caller go through redact().

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@saucam

saucam commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator Author

@highflame-oracle review

The open-questions line still said spawns default to Claude, contradicting
the line #339 updated earlier in the same doc.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@saucam
saucam merged commit cafce89 into main Sep 27, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add a top-level defaultProvider config key — the daemon default is hardcoded to "claude"

2 participants