Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
38 changes: 37 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,6 +75,42 @@ to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

### Fixed

- **The context window is no longer guessed from the model id.**
codeoid inferred every window from a substring table over model ids, and that table was wrong the moment a model shipped: `claude-opus-5-5` inferred to the 200k fallback while every turn result stated 1,000,000.
The percent-of-window display and the history seed sized for a fork or a provider switch were both built on that guess.

The backends already state it, and the daemon now uses what they state.
Audited across every backend codeoid drives, because they do not agree on how they publish it:

| Backend | States | Where |
| --- | --- | --- |
| claude | per turn | `result.modelUsage[model].contextWindow` |
| codex | per turn | `thread/tokenUsage/updated` → `tokenUsage.modelContextWindow` |
| pi | per turn, and on its catalog | `get_session_stats` → `contextUsage.contextWindow`; model objects on `get_available_models` |
| qwen | on its catalog | `contextWindowSize` per model |
| gemini · openai · acp | nothing | — |

Three of those were already arriving and being discarded: codex's on a notification the provider handled while reading two of its three fields, pi's on a stats call made every turn, and qwen's in a catalog projection that dropped it (and, in API-key mode, a catalog union that dropped it again whenever the gateway listed the same model).

One resolver now answers for every consumer: the display, the occupancy caps, auto-rotate, the rotate message, and the seed.
It uses the window this session's backend stated for the model it is running, then one stated in the same scope or published on the provider's catalog, then a per-provider floor.
The stated window is dropped whenever the model or provider changes, so a switch sizes the incoming backend's seed for that backend.
A fork on the same model as its parent inherits the parent's stated window.

What a turn states is remembered per account, project and workdir, not daemon-wide.
The Claude CLI derives the number from settings a workdir can override (`CLAUDE_CODE_MAX_CONTEXT_TOKENS` is unclamped), so one workdir must not size another tenant's seeds.
A window published on a catalog is read from the catalog, so it goes away when the backend stops publishing it.

Auto-rotate used a fixed 1M window rather than the table, so for any model below about 970k its 0.97 hard ceiling could never fire before the backend's own limit — a 200k Haiku or a 272k codex session never rotated.
It now rotates against the window of the model actually running.
The TUI read neither the table nor the stated window; it now shows the same number as the web UI.
With the memory engine off, the window was never reported to clients at all; it now always is.

The static tables stay as the floor for what no backend has stated: before a first turn, and on backends that publish nothing.

Also fixed at the source: the Claude provider took the turn's model from `Object.keys(modelUsage)[0]`, but that map is keyed by every model the turn touched in insertion order, and a Haiku side-call routinely sits first.
It now uses the model the SDK names on `init`, and follows `model_fallback`, so a turn the fallback model served is labelled with that model and its window.

- **`opus` ran Opus 5, months after Opus 5.5 shipped — and a 1M model was measured against a 200k window.**
Bumps `@anthropic-ai/claude-agent-sdk` 0.3.258 → 0.3.281.
The SDK version tracks the bundled Claude Code CLI 1:1, and the model list comes from that CLI at runtime, so the dependency bump is what surfaces a new model at all: on 0.3.258 the backend reported the `opus` alias as "Opus 5", on 0.3.281 it reports "Opus 5.5".
Expand All @@ -89,7 +125,7 @@ to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

The catalog and the window table are now cross-checked in tests: every catalog entry's declared `contextWindow` must equal what `contextWindowForModel` derives from its id and its alias. That is the assertion that would have caught this, and it fails if either table moves without the other.

The deeper fix is separate and tracked: the SDK already reports `modelUsage[model].contextWindow` per turn and codeoid discards it, which is why a hardcoded table existed to go stale. Threading the provider-reported number through — with the static tables demoted to a pre-first-turn bootstrap — is a follow-up, since the window is unknown until a turn completes and some backends report none.
The deeper fix — taking the window from the backend instead of the table — is the entry above.

- **A collaboration's orchestrator could not read the artifacts its own constitution told it to synthesize** (#338).
The compiled goal pack instructs it to "read each role's artifact from the blackboard, merge, and show disagreement", while its default profile read `spec` and `findings` only.
Expand Down
17 changes: 13 additions & 4 deletions packages/protocol/src/types.ts
Original file line number Diff line number Diff line change
Expand Up @@ -474,10 +474,12 @@ export interface SessionUsage {
/** Most recent turn's cache-read ratio (cache_read / total_input). */
lastTurnCacheHitRate?: number;
/**
* Resolved model's context window in tokens — the denominator for
* ctx-occupancy displays. Derived from `SessionInfo.model` via the
* daemon's per-model catalog (`contextWindowForModel`). Switching
* models mid-session updates this on the next info_update broadcast.
* Context window in tokens of the model this session is running — the
* denominator for ctx-occupancy displays, and the one the daemon's own
* auto-rotate uses. The window the backend reported on its last turn when
* there is one, else what the provider publishes for the model, else a
* per-provider estimate. Switching model or provider drops the reported
* value, so this follows the switch on the next info_update.
*
* Optional for back-compat with daemons that pre-date this field;
* frontends should fall back to a conservative constant (200k) or
Expand Down Expand Up @@ -1686,6 +1688,13 @@ export interface ModelInfo {
description?: string;
/** True for the backend's recommended default. */
isDefault?: boolean;
/**
* Context window in tokens, when the backend publishes it on its catalog
* (qwen-code's `contextWindowSize`, pi's model objects). Known before any
* turn runs. Absent where the catalog carries none — Claude and codex report
* the window per turn instead, and gemini, openai and acp not at all.
*/
contextWindow?: number;
}

/**
Expand Down
5 changes: 2 additions & 3 deletions src/daemon/providers/acp/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,7 @@ import type {
ToolApprovalFn,
TurnOpts,
TurnRun,
CatalogEntry,
} from "../interface.js";
import { renderHistorySeed, type CanonicalTurn, type HistorySeedResult } from "../canonical.js";
import { buildGeminiCliEnv } from "../env.js";
Expand Down Expand Up @@ -71,9 +72,7 @@ export interface GeminiAcpProviderInit {
/** Cross-backend MCP registry — external servers mount on session/new
* (gemini-cli owns its client); approval flows through canUseTool. */
mcpRegistry?: McpRegistry;
onModels?: (
models: ReadonlyArray<{ value: string; displayName: string; description?: string }>,
) => void;
onModels?: (models: ReadonlyArray<CatalogEntry>) => void;
}

export class GeminiAcpProvider implements SessionProvider {
Expand Down
61 changes: 53 additions & 8 deletions src/daemon/providers/claude/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ import { FLEET_TOOL_NAMES } from "../../fleet.js";
import { rewriteBashToolInput } from "../../compress/index.js";
import type { CodeoidConfig } from "../../../config.js";
import type { AuthContext } from "../../../protocol/types.js";
import type { SessionProvider, ModelInfo, NormalizedTurnResult, ProviderEvent, SessionScopedEvent, TurnOpts, TurnRun } from "../interface.js";
import type { SessionProvider, ModelInfo, NormalizedTurnResult, ProviderEvent, SessionScopedEvent, TurnOpts, TurnRun, CatalogEntry } from "../interface.js";
import { renderHistorySeed, type CanonicalTurn, type HistorySeedResult } from "../canonical.js";
import { buildSubprocessEnv, withGatewayCredential } from "../env.js";
import type { LLMCallUsage } from "../../context-math.js";
Expand Down Expand Up @@ -125,7 +125,7 @@ export interface ClaudeProviderInit {
config?: CodeoidConfig;
compressionRegistry?: CompressionRegistry;
/** Called once per session with the live model catalog. */
onModels?: (models: ReadonlyArray<{ value: string; displayName: string; description?: string }>) => void;
onModels?: (models: ReadonlyArray<CatalogEntry>) => void;
/**
* Called when the backing Claude Code session is missing (i.e. the SDK
* throws "No conversation found with session ID"). Session must enqueue
Expand Down Expand Up @@ -166,7 +166,7 @@ export class ClaudeProvider implements SessionProvider {
#lastPushedContent: string | null = null;
/** Cross-message carry-over for the translator (see TranslateState) — holds
* the `<local-command-stderr>` that explains a zero-turn "success". */
#translateState: TranslateState = { lastLocalCommandStderr: null };
#translateState: TranslateState = { lastLocalCommandStderr: null, primaryModel: null };
/** Skill commands with an approval prompt in flight — prevents a repeated
* turn from raising a duplicate prompt for the same command (#233). */
#pendingSkillApprovals = new Set<string>();
Expand Down Expand Up @@ -1118,6 +1118,19 @@ export class ClaudeProvider implements SessionProvider {
export interface TranslateState {
/** Last `<local-command-stderr>` seen since the previous result. */
lastLocalCommandStderr: string | null;
/**
* The PRIMARY model id for the turn, as the SDK named it on its `init`
* message (e.g. `claude-opus-5-5`).
*
* Load-bearing for two things that used to guess. `result.modelUsage` is a
* map keyed by every model the turn touched, and side-calls (title
* generation, summaries) put Haiku in there too — often FIRST, since the
* order is insertion, not significance. So `Object.keys(modelUsage)[0]`
* labelled an Opus turn as Haiku, and reading a context window off that
* entry would report 200k for a 1M turn: the very bug being fixed, in a new
* place. The SDK already says which model is primary; ask it instead.
*/
primaryModel: string | null;
}

/**
Expand All @@ -1130,7 +1143,7 @@ export function translateSDKMessage(
msg: SDKMessage,
emit: (event: ProviderEvent) => void,
providerId: string,
state: TranslateState = { lastLocalCommandStderr: null },
state: TranslateState = { lastLocalCommandStderr: null, primaryModel: null },
/** Session-scoped emitter — background-task events go HERE, never through
* `emit`, because they can fire between turns when the turn queue has no
* reader and would drop them unread. */
Expand Down Expand Up @@ -1235,10 +1248,26 @@ export function translateSDKMessage(
cache_read_input_tokens?: number;
cache_creation_input_tokens?: number;
};
modelUsage?: Record<string, { inputTokens?: number; outputTokens?: number }>;
modelUsage?: Record<
string,
{
inputTokens?: number;
outputTokens?: number;
contextWindow?: number;
}
>;
};
// Derive the model from the first key in modelUsage (most-used model this turn).
const model = Object.keys(r.modelUsage ?? {})[0] ?? "unknown";
// The primary model the SDK named at `init`, not the first key in
// modelUsage: that map is keyed by every model the turn touched and
// ordered by insertion, so a Haiku side-call routinely sits in front
// of the model that did the work.
const usageKeys = Object.keys(r.modelUsage ?? {});
const model = state.primaryModel ?? usageKeys[0] ?? "unknown";
// Limits as the BACKEND states them for that model. Authoritative:
// codeoid otherwise guesses the window from a substring table that
// goes stale every time a model ships (see context-windows.ts).
const primaryUsage = r.modelUsage?.[model];
const reportedWindow = primaryUsage?.contextWindow;
// A zero-turn run means the assistant never ran — the prompt was
// consumed and discarded (blocked slash-command expansion, rejected
// input). The SDK reports this as `subtype: "success", is_error: false`,
Expand All @@ -1261,6 +1290,7 @@ export function translateSDKMessage(
cacheCreationTokens: r.usage?.cache_creation_input_tokens ?? 0,
totalCostUsd: r.total_cost_usd ?? 0,
durationMs: r.duration_ms ?? 0,
...(typeof reportedWindow === "number" && reportedWindow > 0 ? { contextWindow: reportedWindow } : {}),
stopReason: r.stop_reason ?? undefined,
// Preserve `undefined` when the SDK didn't report — only force true.
isError: zeroTurn ? true : r.is_error,
Expand All @@ -1276,7 +1306,14 @@ export function translateSDKMessage(
case "system": {
const subtype = (msg as { subtype?: string }).subtype;
if (subtype === "init") {
const init = msg as { mcp_servers?: { name: string; status: string }[]; tools?: string[] };
const init = msg as {
mcp_servers?: { name: string; status: string }[];
tools?: string[];
model?: string;
};
// Remember which model is actually driving this turn — see
// TranslateState.primaryModel for why the result map can't tell us.
if (init.model) state.primaryModel = init.model;
const servers: Record<string, string> = {};
const tools: Record<string, string[]> = {};
for (const s of init.mcp_servers ?? []) {
Expand All @@ -1293,6 +1330,14 @@ export function translateSDKMessage(
tools[server].push(t);
}
emit({ type: "mcp_init", servers, tools });
} else if (subtype === "model_fallback" || subtype === "model_refusal_fallback") {
// The CLI re-dispatched this turn to the configured fallbackModel
// (primary overloaded / refused). The turn is now that model's, so
// its label and its window must be too — otherwise a turn Haiku
// served is recorded as Opus with Opus's 1M. The next turn's `init`
// names the primary again, which restores it.
const fb = (msg as { fallback_model?: string }).fallback_model;
if (fb) state.primaryModel = fb;
} else if (subtype === "api_retry") {
const r = msg as { attempt?: number; retry_delay_ms?: number; error_status?: number | null };
emit({ type: "api_retry", attempt: r.attempt, retryDelayMs: r.retry_delay_ms, errorStatus: r.error_status });
Expand Down
29 changes: 24 additions & 5 deletions src/daemon/providers/codex/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,7 @@ import type {
TurnOpts,
TurnRun,
UiRequestFn,
CatalogEntry,
} from "../interface.js";
import { renderHistorySeed, type CanonicalTurn, type HistorySeedResult } from "../canonical.js";
import { buildCodexEnv } from "../env.js";
Expand Down Expand Up @@ -106,9 +107,7 @@ export interface CodexProviderInit {
/** Cross-backend MCP registry — external servers mount natively via `-c
* mcp_servers.*` (codex owns its client); approval flows through canUseTool. */
mcpRegistry?: McpRegistry;
onModels?: (
models: ReadonlyArray<{ value: string; displayName: string; description?: string }>,
) => void;
onModels?: (models: ReadonlyArray<CatalogEntry>) => void;
}

/** Item types that surface as tool records (vs text/reasoning streams). */
Expand Down Expand Up @@ -272,6 +271,8 @@ export class CodexProvider implements SessionProvider {
* counts. Captured here and folded into turn_done.
*/
#lastTokenUsage: CodexTokenUsage | null = null;
/** Window codex reported for the model driving this thread, if it has. */
#modelContextWindow: number | null = null;
/** item id → {name, input} for items already announced via tool_start. */
#announcedItems = new Map<string, { name: string; input: Record<string, unknown> }>();
/**
Expand Down Expand Up @@ -430,6 +431,11 @@ export class CodexProvider implements SessionProvider {
this.#turnModel = opts.model ?? "codex";
this.#announcedItems.clear();
this.#lastTokenUsage = null;
// Per turn, like the usage beside it. This provider instance outlives a
// `/model` switch or a pipeline overrideModel, so a turn interrupted before
// its first token-usage notification would otherwise report the PREVIOUS
// model's window under the new model's id — and that pair was persisted.
this.#modelContextWindow = null;
this.#hasQueried = true;

void this.#startTurn(opts).catch((err: unknown) => {
Expand Down Expand Up @@ -674,15 +680,28 @@ export class CodexProvider implements SessionProvider {
cacheCreationTokens: 0,
totalCostUsd: 0,
durationMs: Date.now() - this.#turnStartedAt,
...(this.#modelContextWindow !== null
? { contextWindow: this.#modelContextWindow }
: {}),
stopReason: (turn?.status as string | undefined) ?? undefined,
};
this.#push({ type: "turn_done", result });
this.#turnQueue?.close();
break;
}
case "thread/tokenUsage/updated": {
const usage = (params.tokenUsage as { last?: CodexTokenUsage } | undefined)?.last;
if (usage) this.#lastTokenUsage = usage;
// `ThreadTokenUsage` has three fields on this wire — `last`, `total`
// and `modelContextWindow` (confirmed in the codex binary's own serde
// descriptors: "struct ThreadTokenUsage with 3 elements"). codeoid read
// only `last` and inferred the window from the model id instead, which
// is the guess this change exists to stop making.
const tu = params.tokenUsage as
| { last?: CodexTokenUsage; modelContextWindow?: number }
| undefined;
if (tu?.last) this.#lastTokenUsage = tu.last;
if (typeof tu?.modelContextWindow === "number" && tu.modelContextWindow > 0) {
this.#modelContextWindow = tu.modelContextWindow;
}
break;
}
case "error": {
Expand Down
Loading
Loading