From 977e8b4d613c628859522194006c77ab0fc980a8 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 08:18:48 +0000 Subject: [PATCH 01/19] task: add workspace callback token renewal Co-Authored-By: Claude Opus 5.5 --- ...-10-04-workspace-callback-token-renewal.md | 163 ++++++++++++++++++ 1 file changed, 163 insertions(+) create mode 100644 tasks/active/2026-10-04-workspace-callback-token-renewal.md diff --git a/tasks/active/2026-10-04-workspace-callback-token-renewal.md b/tasks/active/2026-10-04-workspace-callback-token-renewal.md new file mode 100644 index 0000000000..182079f324 --- /dev/null +++ b/tasks/active/2026-10-04-workspace-callback-token-renewal.md @@ -0,0 +1,163 @@ +# Secure renewal of workspace callback tokens for long-lived sessions + +SAM task `01M42YQA8QPJQBW48KTDQFAHDE` (parent `01M42WJSH7238RWH5ZH7TSSFZG`), branch +`sam/fix-secure-renewal-workspace-qfahde`. + +## Problem + +Workspace-scoped VM callback JWTs (`signCallbackToken`, `apps/api/src/services/jwt.ts`, +default 24h via `CALLBACK_TOKEN_EXPIRY_MS`) are minted only when a workspace is created or +restored. The node heartbeat renews only the **node** token (`node-lifecycle.ts` heartbeat → +`refreshedToken` → `health.go:setCallbackToken`). A workspace awake more than 24h therefore +gets `401 Invalid or expired callback token` on every workspace-scoped callback. + +### Production evidence (read-only, Workers Logs, 100% head sampling) + +- `POST /api/workspaces/01M3Z4CGTEWNBP3VSTFQK90XJ1/session-snapshot/prepare` → 401 repeatedly + 2026-10-03 22:37Z … 2026-10-04 05:57Z (83 in the rolling 24h per the brief; the sleep could + never complete). +- `POST /api/workspaces/01M3XX0HCA4H5XHXB9ZF3N0867/git-token` → 401 at 2026-10-03 12:31:57Z, + about 24h after that workspace started (git credential fill failed). +- `POST /api/projects/01KHRJGANBBWGDY1NZ0KVF0D4J/workspace-resource-history` → 401 several + times 10-03 22:24Z … 10-04 05:11Z (resource-history spool deleted as "permanent"). +- `POST …/messages` → **no 401 observed** 10-01 … 10-04: both long-lived sessions produced no + agent output after their 24h mark (last message 200s for 01M3XX0H… at 10-02 20:56Z). Message + loss is proven by code below, not by production observation. + +## Research findings + +### Token custody (VM agent) + +- Canonical per-workspace store: `WorkspaceRuntime.CallbackToken` (`internal/server/server.go`), + persisted encrypted in SQLite (`workspace_runtime_persistence.go`) and hydrated on restart + (`workspace_routing.go` `upsertWorkspaceRuntime`). Token-only updates are **not** persisted today + (`metadataChanged` is not set when only the token changes). +- Writers today: create-workspace, `UpdateAfterBootstrap`, a hibernate request body + (`session_snapshot.go` `sessionSnapshotHandlerInput`, supported since 2026-07-11), and the + restore body. The API only sends a body token on cf-container restore + (`vm-agent-container.ts`), never on VM hibernate (`node-agent-session-snapshots.ts`). +- Consumers (24 call sites, full table in the session notes). Three custody classes: + 1. **Live reads** of `runtime.CallbackToken` at call time: git-token (helper and agent start), + runtime-assets, task status callbacks, credential sync, eviction delivery (retried every + heartbeat forever), resource history (401 ⇒ spool deleted), provisioning/recovery. These heal + automatically once the runtime token is renewed. + 2. **Copies that outlive the request**: ACP `SessionHost` `h.config.CallbackToken` (agent-key, + agent-settings, activity, usage, interactions/elicitation, URL completion); message reporter + `authToken`; snapshot capture input (prepare/progress/complete/failure/artifacts); publish jobs. + 3. **Baked into the agent subprocess** at process start: platform AI proxy credential + (`ANTHROPIC_AUTH_TOKEN`/`OPENAI_API_KEY`, `{wstoken}` base URLs, Codex `config.toml`) and the + codex refresh URL. Cannot be rotated without restarting the agent process. +- `messagereport` treats 401 as **terminal**: `markTerminalPersistenceFailure` clears the session + outbox and latches `terminalPersistenceFailure`; `Enqueue` then silently drops every later message + (`sender.go:isTerminalBatchResponse`, `reporter.go`). So a chat awake >24h that produces output + loses that output permanently until the agent restarts (code-proven; reproduced in tests). +- The workspace token is visible inside the devcontainer (agent env for SAM-proxy mode, Codex + config). The node token is not in VM devcontainers (cf-container: the agent inherits the node + token from the container env). + +### Control plane + +- `verifyWorkspaceCallbackAuth` is stateless (signature/scope/claim). Revocation of workspace + callbacks is enforced per route through D1 state (`assertWorkspaceAcceptsCallback`: workspace + status `creating|running|recovery`, node non-terminal, else 410). +- Message writes are idempotent server-side (`project-data/messages.ts`: dedupe by message id and + user-content), so resending a batch that got 401 cannot duplicate rows. A 401 is raised before the + body is read, so nothing was persisted. +- Instant (cf-container) stale-callback guard (`routes/_stale-callback-guard.ts`) compares the + token `iat` to the row `updated_at`. A renewal that refreshes `iat` would hide a superseded + container generation. Renewed tokens must carry the original generation issue time. +- Placement binds `workspaces.node_id` with `nodes.user_id = workspace user` and never moves a bound + workspace (attach requires `node_id IS NULL`; evicted restart keeps the node). +- Existing dual-credential precedent: snapshot upload relay requires the workspace bearer plus the + relay node's node-scoped bearer (`session-snapshot-upload-relay.ts` + `verifySessionSnapshotRelayAuthorization`, headers `X-SAM-Relay-Node-ID` / + `X-SAM-Relay-Authorization`). + +### Design (no new trust boundary) + +1. **API push on hibernate (heals running old agents, no binary rollout needed):** + `hibernateAgentSessionOnNode` mints a fresh workspace token only when D1 confirms the workspace + is active on the target node, and sends it as `workspaceCallbackToken` over the existing + node-management channel (same channel that delivers the token at create/restore). Agents since + 2026-07-11 already store it before capture. +2. **Proactive dual-proof renewal (new agents):** `POST /api/workspaces/:id/callback-token/renew` + requires the current, unexpired workspace token (`Authorization`) **and** the hosting node's + node token (`X-SAM-Node-ID` / `X-SAM-Node-Authorization`). It renews only if the workspace is + active, bound to that node, owned by the node's user, the node is non-terminal, and (Instant) the + token generation is not superseded. A renewed token preserves the generation issue time (`gen` + claim) so the stale-callback guard keeps working. Not-yet-due tokens are not re-minted. + - A node token alone cannot obtain a workspace token (no widened node authority). + - A workspace token leaked from a VM devcontainer cannot renew itself (needs the node token). + - Expired tokens are never renewed (no expiry bypass); recovery is the control-plane push. +3. **Agent:** after each successful heartbeat, renew due tokens (past + `WORKSPACE_CALLBACK_TOKEN_REFRESH_RATIO` of lifetime, default 0.5) with bounded backoff, latch + definitive rejections, compare-and-swap the runtime token, persist it, and propagate every token + change (renewal or control-plane push) to the message reporter and ACP session hosts. +4. **Message reporter:** a 401 is no longer terminal-and-destructive. If a newer token exists, retry + with it; otherwise keep the outbox, send nothing, and resume when a new token arrives. Bounded by + the existing outbox cap and a new configurable park budget, after which the old terminal + behaviour applies (rule 54.13). +5. **ACP SessionHost:** read the callback token through a lock-free accessor that renewal updates + (rule 46: nothing reachable from the ACP notification goroutine may take `mu`). + +## Implementation checklist + +### API +- [ ] `jwt.ts`: renewal signing preserves generation (`gen` claim); payload exposes generation; + stale-callback guard reads generation before `iat` +- [ ] New service `workspace-callback-token-renewal.ts`: dual-proof verification, D1 binding checks, + due check, Instant superseded check, mint; hibernate-delivery mint helper (active + bound) +- [ ] New callback route file `routes/workspaces/callback-token-renewal.ts` mounted in + `routes/workspaces/index.ts`; `Cache-Control: no-store`; designed 401/403/410; IDs-only logs +- [ ] `node-agent-session-snapshots.ts`: include fresh token on hibernate when bound + active +- [ ] Env vars: `WORKSPACE_CALLBACK_TOKEN_RENEWAL_*` documented (`env.ts`, `.env.example`, docs) + +### VM agent +- [ ] Config: renewal ratio, retry initial/max, request timeout; heartbeat parse unaffected +- [ ] `workspace_callback_token_renewal.go`: due selection with injected clock, request with both + tokens, response classification, bounded backoff, rejection latch, CAS apply +- [ ] Hook renewal after successful heartbeat (TryLock, like ready/eviction retries) +- [ ] `upsertWorkspaceRuntime`: persist token changes and propagate to reporter + session hosts +- [ ] `acp.SessionHost`: lock-free current-token accessor + `SetCallbackToken`; replace reads +- [ ] `messagereport`: stale-token retry, park-on-401, resume on new token, bounded park budget + +### Tests +- [ ] Workers test through the real route: renew success (gen preserved, scope workspace, new exp), + expired token denied, node token as workspace proof denied, workspace token as node proof + denied, foreign node denied, workspace on other node (moved) denied, user mismatch denied, + deleted/stopped/evicted workspace 410, terminal node 410, claim/path mismatch, not-due no mint, + Instant superseded denied, concurrent renewals +- [ ] Unit: hibernate push includes token only when bound + active; stale guard uses `gen` +- [ ] Go: clock crosses threshold/24h; expired not sent; rejection latch; backoff; CAS vs concurrent + push; heartbeat (node) vs workspace token separation; propagation to reporter + session host; + persistence across restart; reporter park/resume/no-duplicate/no-loss/budget; real HTTP path +- [ ] `go test -race` for the touched packages + +### Docs / rollout +- [ ] Public docs: security architecture (callback token lifetime + renewal), env reference +- [ ] Rollout note: API push heals old agents for snapshot calls only; other consumers need the new + agent (new nodes); no hot replacement; AI proxy env token residual risk → SAM Idea + +## Acceptance criteria + +- [ ] A VM workspace awake past 24h keeps working: snapshot prepare/progress/complete, git-token, + runtime assets, messages, ACP activity/usage/interactions (new agent) +- [ ] Old agents: hibernate/snapshot callbacks succeed after 24h via the API push alone +- [ ] Renewal never mints for: expired tokens, node-only callers, foreign nodes, moved/deleted/ + stopped workspaces, terminal nodes, superseded Instant generations +- [ ] No message loss or duplication across a token rotation; 401 handling is bounded +- [ ] No tokens in logs; responses carrying tokens are `no-store` +- [ ] Independent security review passes; parent informed of the design + +## Staging + +User permitted skipping staging for this wave. Reason: the failure needs a workspace awake more +than 24h (or a VM with an injected clock), and the cross-boundary contract is fully exercised by +workers tests through the real route plus Go tests against an HTTP control plane. Substitute: +deterministic clock-injected Go tests, Miniflare workers tests through the real auth boundary, +`-race` runs, and post-deploy production log checks. + +## References + +- `.claude/rules/28-credential-resolution-fallback-tests.md`, `.claude/rules/34` (callback auth), + `packages/vm-agent/.claude/rules/54-vm-agent-rollout-compatibility.md`, rule 46, rule 62, rule 73 From ef688adf2ecfb572c4575e4c23a504a0ba4f2341 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 08:38:01 +0000 Subject: [PATCH 02/19] feat(api): renew workspace callback tokens with dual node+workspace proof Workspace callback JWTs expire after CALLBACK_TOKEN_EXPIRY_MS (24h) but a workspace can stay awake longer; every snapshot, git-token and resource-history callback then failed with 401. - POST /api/workspaces/:id/callback-token/renew renews the current, unexpired workspace token only with the hosting node's token as a second proof, only while D1 binds the workspace to that node (same owner) and the workspace and node are active, and only past the refresh ratio. Renewed tokens keep the generation issue time (gen_iat) so the Instant stale-callback guard still detects superseded containers. - hibernateAgentSessionOnNode delivers a fresh workspace token over the node-management channel when the workspace is bound and active; VM agents already store it, so running nodes heal without a binary rollout. Co-Authored-By: Claude Opus 5.5 --- apps/api/src/routes/_stale-callback-guard.ts | 21 +- .../workspaces/callback-token-renewal.ts | 60 +++ apps/api/src/routes/workspaces/index.ts | 2 + .../api/src/services/callback-token-claims.ts | 48 ++ apps/api/src/services/jwt.ts | 16 +- .../services/node-agent-session-snapshots.ts | 21 +- .../workspace-callback-token-renewal.ts | 227 ++++++++++ .../workspace-deletion-callback-signal.ts | 1 + .../unit/routes/stale-callback-guard.test.ts | 41 ++ .../workspace-callback-token-renewal.test.ts | 425 ++++++++++++++++++ 10 files changed, 848 insertions(+), 14 deletions(-) create mode 100644 apps/api/src/routes/workspaces/callback-token-renewal.ts create mode 100644 apps/api/src/services/callback-token-claims.ts create mode 100644 apps/api/src/services/workspace-callback-token-renewal.ts create mode 100644 apps/api/tests/workers/workspace-callback-token-renewal.test.ts diff --git a/apps/api/src/routes/_stale-callback-guard.ts b/apps/api/src/routes/_stale-callback-guard.ts index 0f5f1be4bc..1669b84105 100644 --- a/apps/api/src/routes/_stale-callback-guard.ts +++ b/apps/api/src/routes/_stale-callback-guard.ts @@ -1,6 +1,5 @@ -import { decodeJwt } from 'jose'; - import type { Env } from '../env'; +import { callbackTokenGenerationIssuedAtSeconds } from '../services/callback-token-claims'; /** * Staleness guard for VM-agent → control-plane DESTRUCTIVE callbacks (S2). @@ -43,20 +42,20 @@ export function getInstantStaleCallbackMarginMs(env: Env): number { } /** - * Read the `iat` (issued-at) claim from an ALREADY-VERIFIED callback token and + * Read the issue time of an ALREADY-VERIFIED callback token's generation and * return it in milliseconds. The token MUST have been verified by * `verifyCallbackToken` first — `decodeJwt` does not verify the signature; it is - * used here only to read a claim `verifyCallbackToken` does not surface. + * used here only to read claims `verifyCallbackToken` does not surface. + * + * A renewed workspace token keeps its chain's first `iat` in the `gen_iat` claim + * (`callbackTokenGenerationIssuedAtSeconds`), so renewal never makes a superseded + * container generation look newer than the recovery that replaced it. * - * Returns null when the token cannot be decoded or has no numeric `iat`. + * Returns null when the token cannot be decoded or has no usable issue time. */ export function callbackTokenIssuedAtMs(token: string): number | null { - try { - const claims = decodeJwt(token); - return typeof claims.iat === 'number' ? claims.iat * 1000 : null; - } catch { - return null; - } + const generationIssuedAtSeconds = callbackTokenGenerationIssuedAtSeconds(token); + return generationIssuedAtSeconds === null ? null : generationIssuedAtSeconds * 1000; } export interface SupersededInstantCallbackInput { diff --git a/apps/api/src/routes/workspaces/callback-token-renewal.ts b/apps/api/src/routes/workspaces/callback-token-renewal.ts new file mode 100644 index 0000000000..860b81da07 --- /dev/null +++ b/apps/api/src/routes/workspaces/callback-token-renewal.ts @@ -0,0 +1,60 @@ +/** + * POST /api/workspaces/:id/callback-token/renew — VM agent callback. + * + * The VM agent calls this (packages/vm-agent/internal/server/workspace_callback_token_renewal.go) + * before a workspace callback token expires, so a workspace that stays awake longer than + * CALLBACK_TOKEN_EXPIRY_MS keeps working. Auth is two callback JWTs, never a session cookie + * (.claude/rules/34): the workspace's current token in `Authorization`, and the hosting + * node's id and node token in the JSON body. See + * `services/workspace-callback-token-renewal.ts` for the binding rules. + * + * `workspacesRoutes` applies no session middleware, so this callback route is safe to + * mount there next to the other workspace callbacks (`/:id/messages`, `/:id/git-token`). + */ +import { Hono } from 'hono'; +import * as v from 'valibot'; + +import type { Env } from '../../env'; +import { extractBearerToken } from '../../lib/auth-helpers'; +import { errors } from '../../middleware/error'; +import { renewWorkspaceCallbackToken } from '../../services/workspace-callback-token-renewal'; + +const MAX_NODE_ID_LENGTH = 128; +const MAX_NODE_TOKEN_LENGTH = 16 * 1024; + +const RenewalRequestSchema = v.object({ + nodeId: v.pipe(v.string(), v.trim(), v.minLength(1), v.maxLength(MAX_NODE_ID_LENGTH)), + nodeToken: v.pipe(v.string(), v.minLength(1), v.maxLength(MAX_NODE_TOKEN_LENGTH)), +}); + +const callbackTokenRenewalRoutes = new Hono<{ Bindings: Env }>(); + +callbackTokenRenewalRoutes.post('/:id/callback-token/renew', async (c) => { + const workspaceId = c.req.param('id'); + const workspaceToken = extractBearerToken(c.req.header('Authorization')); + + // The body carries a credential, so validation failures must never echo it back + // (jsonValidator interpolates offending values into its 400 message). + let raw: unknown; + try { + raw = await c.req.json(); + } catch { + throw errors.badRequest('Invalid callback token renewal request'); + } + const parsed = v.safeParse(RenewalRequestSchema, raw); + if (!parsed.success) { + throw errors.badRequest('Invalid callback token renewal request'); + } + + const result = await renewWorkspaceCallbackToken(c.env, { + workspaceId, + workspaceToken, + nodeId: parsed.output.nodeId, + nodeToken: parsed.output.nodeToken, + }); + // The response can carry a credential; no intermediary may store it. + c.header('Cache-Control', 'no-store'); + return c.json(result); +}); + +export { callbackTokenRenewalRoutes }; diff --git a/apps/api/src/routes/workspaces/index.ts b/apps/api/src/routes/workspaces/index.ts index 98e0f6b8a4..2ac54ecce4 100644 --- a/apps/api/src/routes/workspaces/index.ts +++ b/apps/api/src/routes/workspaces/index.ts @@ -3,6 +3,7 @@ import { Hono } from 'hono'; import type { Env } from '../../env'; import { agentSessionSuspendResumeRoutes } from './agent-session-suspend-resume'; import { agentSessionRoutes } from './agent-sessions'; +import { callbackTokenRenewalRoutes } from './callback-token-renewal'; import { crudRoutes } from './crud'; import { lifecycleRoutes } from './lifecycle'; import { localForwardRoutes } from './local-forward'; @@ -17,5 +18,6 @@ workspacesRoutes.route('/', agentSessionRoutes); workspacesRoutes.route('/', agentSessionSuspendResumeRoutes); workspacesRoutes.route('/', runtimeRoutes); workspacesRoutes.route('/', sessionSnapshotRoutes); +workspacesRoutes.route('/', callbackTokenRenewalRoutes); export { workspacesRoutes }; diff --git a/apps/api/src/services/callback-token-claims.ts b/apps/api/src/services/callback-token-claims.ts new file mode 100644 index 0000000000..83cae79c36 --- /dev/null +++ b/apps/api/src/services/callback-token-claims.ts @@ -0,0 +1,48 @@ +/** + * Claim readers for callback JWTs that `verifyCallbackToken` does not surface. + * + * Kept apart from `jwt.ts` (signing and verification) so the guards that read these + * claims do not depend on the signer module. Every function here uses `decodeJwt`, + * which does NOT verify the signature: callers must verify the token first. + */ +import { decodeJwt } from 'jose'; + +/** + * Claim carried by a RENEWED workspace callback token: the `iat` (seconds) of the first + * token in its renewal chain. Renewal must not move a token's generation forward, because + * the Instant stale-callback guard compares the generation's issue time with the most + * recent recovery (see `routes/_stale-callback-guard.ts`). First issuance omits the claim; + * the token's own `iat` is then the generation. + */ +export const CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM = 'gen_iat'; + +function positiveIntegerClaim(value: unknown): number | null { + return typeof value === 'number' && Number.isSafeInteger(value) && value > 0 ? value : null; +} + +/** + * Generation issue time (seconds) of an already-verified callback token: the `gen_iat` + * claim a renewal preserved, else the token's own `iat`. Null when neither is a positive + * integer. + */ +export function callbackTokenGenerationIssuedAtSeconds(token: string): number | null { + try { + const claims = decodeJwt(token); + return ( + positiveIntegerClaim(claims[CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM]) ?? + positiveIntegerClaim(claims.iat) + ); + } catch { + return null; + } +} + +/** `exp` (ms) of an already-verified callback token; null if absent. */ +export function callbackTokenExpiresAtMs(token: string): number | null { + try { + const exp = positiveIntegerClaim(decodeJwt(token).exp); + return exp === null ? null : exp * 1000; + } catch { + return null; + } +} diff --git a/apps/api/src/services/jwt.ts b/apps/api/src/services/jwt.ts index b70ee472b6..34d35cf337 100644 --- a/apps/api/src/services/jwt.ts +++ b/apps/api/src/services/jwt.ts @@ -14,6 +14,7 @@ import { import type { Env } from '../env'; import { AppError } from '../middleware/error'; +import { CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM } from './callback-token-claims'; // Key ID format: key-YYYY-MM (rotates monthly) const KEY_ID = `key-${new Date().getFullYear()}-${String(new Date().getMonth() + 1).padStart(2, '0')}`; @@ -102,6 +103,11 @@ export async function signTerminalToken( }; } +export interface SignCallbackTokenOptions { + /** Generation issue time (seconds) to preserve when renewing an existing token. */ + generationIssuedAtSeconds?: number; +} + /** * Sign a workspace-scoped callback token for VM-to-API authentication. * Used by VM agent to call back to control plane for workspace-specific operations @@ -110,16 +116,24 @@ export async function signTerminalToken( * The `scope: 'workspace'` claim restricts this token to the specific workspace. * Node-scoped tokens cannot be used for workspace-scoped endpoints. */ -export async function signCallbackToken(workspaceId: string, env: Env): Promise { +export async function signCallbackToken( + workspaceId: string, + env: Env, + options: SignCallbackTokenOptions = {} +): Promise { const privateKey = await importPKCS8(env.JWT_PRIVATE_KEY, 'RS256'); const expiry = getCallbackTokenExpiry(env); const expiresAt = new Date(Date.now() + expiry); const issuer = getIssuer(env); + const generationIssuedAt = options.generationIssuedAtSeconds; const token = await new SignJWT({ workspace: workspaceId, type: 'callback', scope: 'workspace', + ...(generationIssuedAt !== undefined + ? { [CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM]: generationIssuedAt } + : {}), }) .setProtectedHeader({ alg: 'RS256', kid: KEY_ID }) .setIssuer(issuer) diff --git a/apps/api/src/services/node-agent-session-snapshots.ts b/apps/api/src/services/node-agent-session-snapshots.ts index 355eedff61..c75e03edcc 100644 --- a/apps/api/src/services/node-agent-session-snapshots.ts +++ b/apps/api/src/services/node-agent-session-snapshots.ts @@ -5,6 +5,7 @@ import { type GuardedNodeAgentMutationOptions, nodeAgentRequest, } from './node-agent'; +import { mintWorkspaceCallbackTokenForNodeDelivery } from './workspace-callback-token-renewal'; export const DEFAULT_SESSION_SNAPSHOT_REQUEST_TIMEOUT_MS = 5 * 60 * 1000; @@ -13,6 +14,12 @@ interface SessionSnapshotRequest { runtime: string; agentType?: string; background?: boolean; + /** + * Fresh workspace-scoped callback token the agent stores before capturing (VM agents + * accept it on hibernate since 2026-07-11). Without it, a workspace awake longer than + * CALLBACK_TOKEN_EXPIRY_MS fails every snapshot callback with 401. + */ + workspaceCallbackToken?: string; } export function getSessionSnapshotRequestTimeoutMs(env: Env): number { @@ -50,7 +57,7 @@ function requestSessionSnapshot( ); } -export function hibernateAgentSessionOnNode( +export async function hibernateAgentSessionOnNode( nodeId: string, workspaceId: string, sessionId: string, @@ -58,7 +65,17 @@ export function hibernateAgentSessionOnNode( userId: string, input: SessionSnapshotRequest ): Promise { - return requestSessionSnapshot('hibernate', nodeId, workspaceId, sessionId, env, userId, input); + // The capture's prepare/progress/complete/failure callbacks authenticate with the + // workspace token the agent holds. Deliver a fresh one over this node-management + // request, exactly as create/restore do, so a long-awake workspace can still sleep. + const workspaceCallbackToken = await mintWorkspaceCallbackTokenForNodeDelivery(env, { + workspaceId, + nodeId, + }); + return requestSessionSnapshot('hibernate', nodeId, workspaceId, sessionId, env, userId, { + ...input, + ...(workspaceCallbackToken ? { workspaceCallbackToken } : {}), + }); } export function restoreAgentSessionOnNode( diff --git a/apps/api/src/services/workspace-callback-token-renewal.ts b/apps/api/src/services/workspace-callback-token-renewal.ts new file mode 100644 index 0000000000..1ed8ed721e --- /dev/null +++ b/apps/api/src/services/workspace-callback-token-renewal.ts @@ -0,0 +1,227 @@ +/** + * Workspace callback token renewal. + * + * Workspace-scoped callback JWTs (`signCallbackToken`) expire after + * CALLBACK_TOKEN_EXPIRY_MS (default 24h), but a workspace can stay awake longer. Two + * existing trusted channels keep the VM agent's copy fresh without widening the + * authority of any credential: + * + * 1. Proof-of-possession renewal (VM agent -> API, `renewWorkspaceCallbackToken`): the + * agent presents the workspace's CURRENT, unexpired token (Authorization header) AND + * its own node-scoped token (JSON body, which Workers Logs never records). Neither + * credential alone is enough, the same dual-credential rule as the snapshot upload + * relay (`verifySessionSnapshotRelayAuthorization`). A node token cannot mint a + * workspace token, a workspace token copied out of a devcontainer cannot renew + * itself, and an expired token is never renewed. + * 2. Control-plane delivery (API -> VM agent, `mintWorkspaceCallbackTokenForNodeDelivery`): + * requests the control plane already sends over the node-management channel carry a + * freshly minted token, exactly as workspace create and cf-container restore do. + * + * Both paths renew only while D1 says the workspace is active and bound to the node, so + * deleting, stopping or reassigning a workspace still ends its callback authority. + */ +import type { Env } from '../env'; +import { log } from '../lib/logger'; +import { AppError, errors } from '../middleware/error'; +import { + assertWorkspaceAcceptsCallback, + assertWorkspaceCallbackIdentityCurrent, + WORKSPACE_CALLBACK_ACTIVE_STATUSES, + type WorkspaceCallbackIdentitySnapshot, +} from '../routes/workspaces/_helpers'; +import { + callbackTokenExpiresAtMs, + callbackTokenGenerationIssuedAtSeconds, +} from './callback-token-claims'; +import { shouldRefreshCallbackToken, signCallbackToken, verifyCallbackToken } from './jwt'; +import { nodeStatusTerminatesCallbacks } from './node-callback-auth'; + +/** + * Error codes for the node credential, distinct from the workspace credential's + * UNAUTHORIZED/FORBIDDEN. The agent retries a node-credential failure after its node + * token refreshes, but stops presenting a workspace token the API rejected. + */ +export const NODE_CALLBACK_UNAUTHORIZED = 'NODE_CALLBACK_UNAUTHORIZED'; +export const NODE_CALLBACK_FORBIDDEN = 'NODE_CALLBACK_FORBIDDEN'; + +interface WorkspaceRenewalBinding extends WorkspaceCallbackIdentitySnapshot { + nodeUserId: string | null; +} + +export type WorkspaceCallbackTokenRenewalResult = + | { renewed: true; token: string; expiresAt: string | null } + | { renewed: false }; + +async function loadWorkspaceRenewalBinding( + env: Env, + workspaceId: string +): Promise { + const row = await env.DATABASE.prepare( + `SELECT w.id AS workspaceId, + w.user_id AS userId, + w.project_id AS projectId, + w.chat_session_id AS chatSessionId, + w.status AS status, + w.node_id AS nodeId, + n.status AS nodeStatus, + n.user_id AS nodeUserId + FROM workspaces w + LEFT JOIN nodes n ON n.id = w.node_id + WHERE w.id = ? + LIMIT 1` + ) + .bind(workspaceId) + .first(); + return row ?? null; +} + +/** + * A workspace's callback authority belongs to the node D1 binds it to. Placement only + * binds a workspace to a node owned by the workspace user, so a mismatch is a foreign + * node and must not receive the workspace's credential. + */ +function workspaceBoundToNode(binding: WorkspaceRenewalBinding, nodeId: string): boolean { + return ( + !!binding.nodeId && + binding.nodeId === nodeId && + !!binding.nodeUserId && + binding.nodeUserId === binding.userId + ); +} + +async function verifyRenewalNodeCredential( + env: Env, + nodeId: string, + nodeToken: string +): Promise { + if (!nodeId || !nodeToken) { + throw new AppError(401, NODE_CALLBACK_UNAUTHORIZED, 'Node callback credential required'); + } + try { + const payload = await verifyCallbackToken(nodeToken, env, { expectedScope: 'node' }); + if (payload.workspace !== nodeId) { + throw new AppError(403, NODE_CALLBACK_FORBIDDEN, 'Node callback token does not match node'); + } + } catch (err) { + if (!(err instanceof AppError)) throw err; + if (err.error === NODE_CALLBACK_FORBIDDEN) throw err; + if (err.statusCode === 401) { + throw new AppError(401, NODE_CALLBACK_UNAUTHORIZED, 'Invalid or expired node callback token'); + } + throw new AppError(403, NODE_CALLBACK_FORBIDDEN, 'Insufficient node token scope'); + } +} + +/** + * Renew a workspace callback token for the agent on the node hosting the workspace. + * + * Returns `{ renewed: false }` while the presented token is younger than the shared + * CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO, so a holder cannot mint early. Throws designed + * 401/403/410 AppErrors; never a 5xx for a rejected credential. + */ +export async function renewWorkspaceCallbackToken( + env: Env, + input: { + workspaceId: string; + workspaceToken: string; + nodeId: string; + nodeToken: string; + } +): Promise { + const { workspaceId, workspaceToken, nodeId } = input; + + // Both credentials are verified before any read, so neither alone can probe state. + const workspacePayload = await verifyCallbackToken(workspaceToken, env, { + expectedScope: 'workspace', + }); + if (workspacePayload.workspace !== workspaceId) { + throw errors.forbidden('Callback token does not match workspace'); + } + await verifyRenewalNodeCredential(env, nodeId, input.nodeToken); + + const binding = await loadWorkspaceRenewalBinding(env, workspaceId); + if (binding && !workspaceBoundToNode(binding, nodeId)) { + log.warn('workspace_callback_token.renewal_rejected', { + workspaceId, + nodeId, + boundNodeId: binding.nodeId, + reason: 'not_bound_to_node', + action: 'rejected', + }); + throw errors.forbidden('Workspace is not hosted on this node'); + } + const active = await assertWorkspaceAcceptsCallback( + env, + binding, + workspaceId, + 'callback_token_renewal' + ); + + if (!shouldRefreshCallbackToken(workspaceToken, env)) { + return { renewed: false }; + } + + const generationIssuedAtSeconds = callbackTokenGenerationIssuedAtSeconds(workspaceToken); + const token = await signCallbackToken(workspaceId, env, { + generationIssuedAtSeconds: generationIssuedAtSeconds ?? undefined, + }); + + // Rule 49: re-read the workspace incarnation at the secret-delivery boundary, so a + // deletion or reassignment that won while minting suppresses the new credential. + await assertWorkspaceCallbackIdentityCurrent(env, active, 'callback_token_renewal'); + + const expiresAtMs = callbackTokenExpiresAtMs(token); + log.info('workspace_callback_token.renewed', { + workspaceId, + nodeId, + generationAgeSeconds: + generationIssuedAtSeconds === null + ? null + : Math.max(0, Math.floor(Date.now() / 1000) - generationIssuedAtSeconds), + }); + return { + renewed: true, + token, + expiresAt: expiresAtMs === null ? null : new Date(expiresAtMs).toISOString(), + }; +} + +/** + * Mint a fresh workspace callback token for a control-plane request that is about to be + * delivered to `nodeId` over the node-management channel. Returns null, and the caller + * sends its request without a token exactly as before, unless D1 binds the workspace to + * that node and the workspace and node are still active. + */ +export async function mintWorkspaceCallbackTokenForNodeDelivery( + env: Env, + input: { workspaceId: string; nodeId: string } +): Promise { + try { + const binding = await loadWorkspaceRenewalBinding(env, input.workspaceId); + const skipReason = !binding + ? 'workspace_missing' + : !workspaceBoundToNode(binding, input.nodeId) + ? 'not_bound_to_node' + : !WORKSPACE_CALLBACK_ACTIVE_STATUSES.has(binding.status) + ? 'workspace_inactive' + : !binding.nodeStatus || nodeStatusTerminatesCallbacks(binding.nodeStatus) + ? 'node_inactive' + : null; + if (skipReason) { + log.info('workspace_callback_token.delivery_skipped', { + workspaceId: input.workspaceId, + nodeId: input.nodeId, + reason: skipReason, + }); + return null; + } + return await signCallbackToken(input.workspaceId, env); + } catch (err) { + log.warn('workspace_callback_token.delivery_mint_failed', { + workspaceId: input.workspaceId, + nodeId: input.nodeId, + error: err instanceof Error ? err.message : String(err), + }); + return null; + } +} diff --git a/apps/api/src/services/workspace-deletion-callback-signal.ts b/apps/api/src/services/workspace-deletion-callback-signal.ts index c722553ff6..8af3892288 100644 --- a/apps/api/src/services/workspace-deletion-callback-signal.ts +++ b/apps/api/src/services/workspace-deletion-callback-signal.ts @@ -14,6 +14,7 @@ export type WorkspaceDeletionCallbackKind = | 'agent_settings' | 'boot_log' | 'bootstrap_token' + | 'callback_token_renewal' | 'compose_image_artifact_complete' | 'compose_image_artifact_init' | 'compose_publish_release' diff --git a/apps/api/tests/unit/routes/stale-callback-guard.test.ts b/apps/api/tests/unit/routes/stale-callback-guard.test.ts index df34ceca13..97e1f83f1e 100644 --- a/apps/api/tests/unit/routes/stale-callback-guard.test.ts +++ b/apps/api/tests/unit/routes/stale-callback-guard.test.ts @@ -1,3 +1,4 @@ +import { exportPKCS8, generateKeyPair } from 'jose'; import { describe, expect, it } from 'vitest'; import type { Env } from '../../../src/env'; @@ -7,6 +8,7 @@ import { getInstantStaleCallbackMarginMs, isSupersededInstantCallback, } from '../../../src/routes/_stale-callback-guard'; +import { signCallbackToken } from '../../../src/services/jwt'; function jwtWith(payload: Record): string { const seg = (obj: unknown) => Buffer.from(JSON.stringify(obj)).toString('base64url'); @@ -31,6 +33,45 @@ describe('callbackTokenIssuedAtMs', () => { expect(callbackTokenIssuedAtMs(jwtWith({ workspace: 'ws-1' }))).toBeNull(); expect(callbackTokenIssuedAtMs(jwtWith({ iat: 'nope' }))).toBeNull(); }); + + it('returns the preserved generation (gen_iat) of a renewed token, not its renewal iat', () => { + expect(callbackTokenIssuedAtMs(jwtWith({ iat: IAT_SECONDS + 86_400, gen_iat: IAT_SECONDS }))).toBe( + IAT_MS + ); + }); + + it('falls back to iat when gen_iat is malformed', () => { + expect(callbackTokenIssuedAtMs(jwtWith({ iat: IAT_SECONDS, gen_iat: 'x' }))).toBe(IAT_MS); + expect(callbackTokenIssuedAtMs(jwtWith({ iat: IAT_SECONDS, gen_iat: -1 }))).toBe(IAT_MS); + expect(callbackTokenIssuedAtMs(jwtWith({ iat: IAT_SECONDS, gen_iat: 1.5 }))).toBe(IAT_MS); + }); + + it('keeps a renewed token from a superseded Instant generation detectable as stale', async () => { + // The renewal route signs with the real signer; reading must agree on the claim name. + const { privateKey } = await generateKeyPair('RS256', { extractable: true }); + const env = { + JWT_PRIVATE_KEY: await exportPKCS8(privateKey), + BASE_DOMAIN: 'example.com', + } as unknown as Env; + const generationIssuedAtSeconds = Math.floor(Date.now() / 1000) - 2 * 3600; + const renewed = await signCallbackToken('ws-instant', env, { generationIssuedAtSeconds }); + const firstIssue = await signCallbackToken('ws-instant', env); + + const recoveredAtMs = generationIssuedAtSeconds * 1000 + MARGIN + 60_000; + const verdict = (token: string) => + isSupersededInstantCallback({ + runtime: 'cf-container', + rowUpdatedAt: iso(recoveredAtMs), + tokenIssuedAtMs: callbackTokenIssuedAtMs(token), + marginMs: MARGIN, + }); + + // Renewed after the recovery, but its generation predates it: still superseded. + expect(callbackTokenIssuedAtMs(renewed)).toBe(generationIssuedAtSeconds * 1000); + expect(verdict(renewed)).toBe(true); + // Control: a generation issued after the recovery is current. + expect(verdict(firstIssue)).toBe(false); + }); }); describe('getInstantStaleCallbackMarginMs', () => { diff --git a/apps/api/tests/workers/workspace-callback-token-renewal.test.ts b/apps/api/tests/workers/workspace-callback-token-renewal.test.ts new file mode 100644 index 0000000000..2de68b252c --- /dev/null +++ b/apps/api/tests/workers/workspace-callback-token-renewal.test.ts @@ -0,0 +1,425 @@ +/** + * Workspace callback token renewal, exercised through the real Worker + * (`SELF.fetch` → index.ts → workspacesRoutes) with real D1 rows and real RS256 tokens. + * + * A workspace callback token is minted at create/restore and lives CALLBACK_TOKEN_EXPIRY_MS + * (24h). Production workspaces awake longer than that got 401 on every snapshot, git-token + * and resource-history callback. These tests drive the renewal route the way the VM agent + * does: an aged-but-valid workspace token plus the hosting node's token. + * + * Tokens are signed with the Worker's real key and the exact claim set `signCallbackToken` + * produces, but with issue/expiry times shifted into the past, which is how the tests cross + * the 50% refresh threshold and the 24h expiry without waiting. + */ +import { env, SELF } from 'cloudflare:test'; +import { decodeJwt, importPKCS8, SignJWT } from 'jose'; +import { afterEach, beforeAll, describe, expect, it, vi } from 'vitest'; + +import type { Env } from '../../src/env'; +import { CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM } from '../../src/services/callback-token-claims'; +import { + signCallbackToken, + signNodeCallbackToken, + verifyCallbackToken, +} from '../../src/services/jwt'; +import { hibernateAgentSessionOnNode } from '../../src/services/node-agent-session-snapshots'; +import { mintWorkspaceCallbackTokenForNodeDelivery } from '../../src/services/workspace-callback-token-renewal'; +import { seedNode, seedUser, seedWorkspace } from './helpers/seed-d1'; + +const testEnv = env as unknown as Env; +const PREFIX = `cbrenew-${Date.now()}`; +const HOUR = 60 * 60; +const DAY = 24 * HOUR; + +const USER_ID = `${PREFIX}-user`; +const OTHER_USER_ID = `${PREFIX}-other-user`; + +const NODE_ID = `${PREFIX}-node`; +const OTHER_NODE_ID = `${PREFIX}-node-2`; +const OTHER_OWNER_NODE_ID = `${PREFIX}-node-other-owner`; +const STOPPED_NODE_ID = `${PREFIX}-node-stopped`; + +const WS_ACTIVE = `${PREFIX}-ws-active`; +const WS_CREATING = `${PREFIX}-ws-creating`; +const WS_ON_OTHER_NODE = `${PREFIX}-ws-other-node`; +const WS_DELETED = `${PREFIX}-ws-deleted`; +const WS_STOPPED = `${PREFIX}-ws-stopped`; +const WS_ON_STOPPED_NODE = `${PREFIX}-ws-stopped-node`; +const WS_OTHER_OWNER_NODE = `${PREFIX}-ws-other-owner-node`; +const WS_MISSING = `${PREFIX}-ws-missing`; + +let nodeToken: string; +let otherNodeToken: string; +let otherOwnerNodeToken: string; +let stoppedNodeToken: string; + +/** + * Same claims as production `signCallbackToken` / `signNodeCallbackToken`, with the issue + * time moved `ageSeconds` into the past. A 24h token aged 13h is past the default 50% + * refresh threshold; aged 25h it is expired. + */ +async function signAgedCallbackToken( + subject: string, + scope: 'workspace' | 'node', + ageSeconds: number, + options: { lifetimeSeconds?: number; generationIssuedAt?: number } = {} +): Promise { + const privateKey = await importPKCS8(testEnv.JWT_PRIVATE_KEY, 'RS256'); + const issuedAt = Math.floor(Date.now() / 1000) - ageSeconds; + return new SignJWT({ + workspace: subject, + type: 'callback', + scope, + ...(options.generationIssuedAt !== undefined + ? { [CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM]: options.generationIssuedAt } + : {}), + }) + .setProtectedHeader({ alg: 'RS256' }) + .setIssuer(`https://api.${testEnv.BASE_DOMAIN}`) + .setSubject(subject) + .setAudience('workspace-callback') + .setIssuedAt(issuedAt) + .setExpirationTime(issuedAt + (options.lifetimeSeconds ?? DAY)) + .sign(privateKey); +} + +function renew( + workspaceId: string, + workspaceToken: string | null, + body: unknown +): Promise { + const headers: Record = { 'Content-Type': 'application/json' }; + if (workspaceToken) headers.Authorization = `Bearer ${workspaceToken}`; + return SELF.fetch( + `https://api.test.example.com/api/workspaces/${workspaceId}/callback-token/renew`, + { + method: 'POST', + headers, + body: typeof body === 'string' ? body : JSON.stringify(body), + } + ); +} + +async function workspaceRow( + workspaceId: string +): Promise<{ status: string; nodeId: string | null; updatedAt: string } | null> { + return testEnv.DATABASE.prepare( + 'SELECT status, node_id AS nodeId, updated_at AS updatedAt FROM workspaces WHERE id = ?' + ) + .bind(workspaceId) + .first(); +} + +async function errorCode(response: Response): Promise { + const body = (await response.json()) as { error?: string }; + return body.error ?? ''; +} + +beforeAll(async () => { + await seedUser(USER_ID); + await seedUser(OTHER_USER_ID); + await seedNode(NODE_ID, USER_ID); + await seedNode(OTHER_NODE_ID, USER_ID); + await seedNode(OTHER_OWNER_NODE_ID, OTHER_USER_ID); + await seedNode(STOPPED_NODE_ID, USER_ID, { status: 'stopped', healthStatus: 'unhealthy' }); + + await seedWorkspace(WS_ACTIVE, NODE_ID, USER_ID, { status: 'running' }); + await seedWorkspace(WS_CREATING, NODE_ID, USER_ID, { status: 'creating' }); + await seedWorkspace(WS_ON_OTHER_NODE, OTHER_NODE_ID, USER_ID, { status: 'running' }); + await seedWorkspace(WS_DELETED, NODE_ID, USER_ID, { status: 'deleted' }); + await seedWorkspace(WS_STOPPED, NODE_ID, USER_ID, { status: 'stopped' }); + await seedWorkspace(WS_ON_STOPPED_NODE, STOPPED_NODE_ID, USER_ID, { status: 'running' }); + // Corrupt binding: placement never puts a workspace on another user's node. + await seedWorkspace(WS_OTHER_OWNER_NODE, OTHER_OWNER_NODE_ID, USER_ID, { status: 'running' }); + + nodeToken = await signNodeCallbackToken(NODE_ID, testEnv); + otherNodeToken = await signNodeCallbackToken(OTHER_NODE_ID, testEnv); + otherOwnerNodeToken = await signNodeCallbackToken(OTHER_OWNER_NODE_ID, testEnv); + stoppedNodeToken = await signNodeCallbackToken(STOPPED_NODE_ID, testEnv); +}); + +describe('POST /api/workspaces/:id/callback-token/renew', () => { + it('renews an aged token for the hosting node and keeps the generation issue time', async () => { + const aged = await signAgedCallbackToken(WS_ACTIVE, 'workspace', 13 * HOUR); + const agedClaims = decodeJwt(aged); + + const response = await renew(WS_ACTIVE, aged, { nodeId: NODE_ID, nodeToken }); + + expect(response.status).toBe(200); + expect(response.headers.get('Cache-Control')).toBe('no-store'); + const body = (await response.json()) as { + renewed: boolean; + token: string; + expiresAt: string; + }; + expect(body.renewed).toBe(true); + const payload = await verifyCallbackToken(body.token, testEnv, { + expectedScope: 'workspace', + }); + expect(payload).toEqual({ workspace: WS_ACTIVE, type: 'callback', scope: 'workspace' }); + const renewedClaims = decodeJwt(body.token); + expect(renewedClaims.sub).toBe(WS_ACTIVE); + expect(renewedClaims.exp).toBeGreaterThan(agedClaims.exp as number); + expect(new Date(body.expiresAt).getTime()).toBe((renewedClaims.exp as number) * 1000); + expect(renewedClaims[CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM]).toBe(agedClaims.iat); + }); + + it('keeps the first generation across a chain of renewals', async () => { + const originalIssuedAt = Math.floor(Date.now() / 1000) - 40 * HOUR; + const agedRenewal = await signAgedCallbackToken(WS_ACTIVE, 'workspace', 13 * HOUR, { + generationIssuedAt: originalIssuedAt, + }); + + const response = await renew(WS_ACTIVE, agedRenewal, { nodeId: NODE_ID, nodeToken }); + + expect(response.status).toBe(200); + const body = (await response.json()) as { token: string }; + expect(decodeJwt(body.token)[CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM]).toBe( + originalIssuedAt + ); + }); + + it('does not mint while the token is younger than the refresh threshold', async () => { + const fresh = await signCallbackToken(WS_ACTIVE, testEnv); + + const response = await renew(WS_ACTIVE, fresh, { nodeId: NODE_ID, nodeToken }); + + expect(response.status).toBe(200); + expect(await response.json()).toEqual({ renewed: false }); + }); + + it('renews a workspace that is still creating', async () => { + const aged = await signAgedCallbackToken(WS_CREATING, 'workspace', 20 * HOUR); + const response = await renew(WS_CREATING, aged, { nodeId: NODE_ID, nodeToken }); + expect(response.status).toBe(200); + expect(((await response.json()) as { renewed: boolean }).renewed).toBe(true); + }); + + it('never renews a token that already expired (24h crossed)', async () => { + const expired = await signAgedCallbackToken(WS_ACTIVE, 'workspace', DAY + HOUR); + + const response = await renew(WS_ACTIVE, expired, { nodeId: NODE_ID, nodeToken }); + + expect(response.status).toBe(401); + expect(await errorCode(response)).toBe('UNAUTHORIZED'); + }); + + it('rejects a node token presented as the workspace credential', async () => { + const response = await renew(WS_ACTIVE, nodeToken, { nodeId: NODE_ID, nodeToken }); + expect(response.status).toBe(403); + expect(await errorCode(response)).toBe('FORBIDDEN'); + }); + + it('rejects a workspace token presented as the node credential', async () => { + const aged = await signAgedCallbackToken(WS_ACTIVE, 'workspace', 13 * HOUR); + const otherWorkspaceToken = await signCallbackToken(WS_ON_OTHER_NODE, testEnv); + + const response = await renew(WS_ACTIVE, aged, { + nodeId: WS_ON_OTHER_NODE, + nodeToken: otherWorkspaceToken, + }); + + expect(response.status).toBe(403); + expect(await errorCode(response)).toBe('NODE_CALLBACK_FORBIDDEN'); + }); + + it('requires the node credential: a workspace token alone cannot renew itself', async () => { + const aged = await signAgedCallbackToken(WS_ACTIVE, 'workspace', 13 * HOUR); + + const missingBody = await renew(WS_ACTIVE, aged, ''); + expect(missingBody.status).toBe(400); + + const emptyNodeToken = await renew(WS_ACTIVE, aged, { nodeId: NODE_ID, nodeToken: '' }); + expect(emptyNodeToken.status).toBe(400); + }); + + it('rejects an expired node token with the node-credential code', async () => { + const aged = await signAgedCallbackToken(WS_ACTIVE, 'workspace', 13 * HOUR); + const expiredNodeToken = await signAgedCallbackToken(NODE_ID, 'node', DAY + HOUR); + + const response = await renew(WS_ACTIVE, aged, { nodeId: NODE_ID, nodeToken: expiredNodeToken }); + + expect(response.status).toBe(401); + expect(await errorCode(response)).toBe('NODE_CALLBACK_UNAUTHORIZED'); + }); + + it('rejects a node token whose claim differs from the named node', async () => { + const aged = await signAgedCallbackToken(WS_ACTIVE, 'workspace', 13 * HOUR); + + const response = await renew(WS_ACTIVE, aged, { nodeId: NODE_ID, nodeToken: otherNodeToken }); + + expect(response.status).toBe(403); + expect(await errorCode(response)).toBe('NODE_CALLBACK_FORBIDDEN'); + }); + + it('refuses a node that does not host the workspace (moved/foreign node), owner control renews', async () => { + const aged = await signAgedCallbackToken(WS_ON_OTHER_NODE, 'workspace', 13 * HOUR); + const before = await workspaceRow(WS_ON_OTHER_NODE); + + const attack = await renew(WS_ON_OTHER_NODE, aged, { nodeId: NODE_ID, nodeToken }); + + expect(attack.status).toBe(403); + expect(await errorCode(attack)).toBe('FORBIDDEN'); + expect(await workspaceRow(WS_ON_OTHER_NODE)).toEqual(before); + + const owner = await renew(WS_ON_OTHER_NODE, aged, { + nodeId: OTHER_NODE_ID, + nodeToken: otherNodeToken, + }); + expect(owner.status).toBe(200); + expect(((await owner.json()) as { renewed: boolean }).renewed).toBe(true); + }); + + it("refuses a node owned by a different user than the workspace's", async () => { + const aged = await signAgedCallbackToken(WS_OTHER_OWNER_NODE, 'workspace', 13 * HOUR); + + const response = await renew(WS_OTHER_OWNER_NODE, aged, { + nodeId: OTHER_OWNER_NODE_ID, + nodeToken: otherOwnerNodeToken, + }); + + expect(response.status).toBe(403); + }); + + it('refuses a token minted for another workspace', async () => { + const tokenForActive = await signAgedCallbackToken(WS_ACTIVE, 'workspace', 13 * HOUR); + + const response = await renew(WS_ON_OTHER_NODE, tokenForActive, { + nodeId: OTHER_NODE_ID, + nodeToken: otherNodeToken, + }); + + expect(response.status).toBe(403); + expect(await errorCode(response)).toBe('FORBIDDEN'); + }); + + it.each([ + ['deleted workspace', WS_DELETED, NODE_ID, () => nodeToken], + ['stopped workspace', WS_STOPPED, NODE_ID, () => nodeToken], + ['workspace on a stopped node', WS_ON_STOPPED_NODE, STOPPED_NODE_ID, () => stoppedNodeToken], + ])('ends renewal for a %s with 410', async (_label, workspaceId, nodeId, token) => { + const aged = await signAgedCallbackToken(workspaceId, 'workspace', 13 * HOUR); + const before = await workspaceRow(workspaceId); + + const response = await renew(workspaceId, aged, { nodeId, nodeToken: token() }); + + expect(response.status).toBe(410); + expect(await errorCode(response)).toBe('GONE'); + expect(await workspaceRow(workspaceId)).toEqual(before); + }); + + it('ends renewal for a workspace row that no longer exists', async () => { + const aged = await signAgedCallbackToken(WS_MISSING, 'workspace', 13 * HOUR); + const response = await renew(WS_MISSING, aged, { nodeId: NODE_ID, nodeToken }); + expect(response.status).toBe(410); + }); + + it('serves concurrent renewals of the same token independently', async () => { + const aged = await signAgedCallbackToken(WS_ACTIVE, 'workspace', 13 * HOUR); + const agedIssuedAt = decodeJwt(aged).iat; + + const responses = await Promise.all( + [0, 1, 2].map(() => renew(WS_ACTIVE, aged, { nodeId: NODE_ID, nodeToken })) + ); + + for (const response of responses) { + expect(response.status).toBe(200); + const body = (await response.json()) as { token: string }; + await expect( + verifyCallbackToken(body.token, testEnv, { expectedScope: 'workspace' }) + ).resolves.toMatchObject({ workspace: WS_ACTIVE }); + expect(decodeJwt(body.token)[CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM]).toBe(agedIssuedAt); + } + }); + + it('lets a workspace callback that failed after 24h succeed with the renewed token', async () => { + const expired = await signAgedCallbackToken(WS_ACTIVE, 'workspace', DAY + HOUR); + const rejected = await SELF.fetch( + `https://api.test.example.com/api/workspaces/${WS_ACTIVE}/runtime`, + { headers: { Authorization: `Bearer ${expired}` } } + ); + expect(rejected.status).toBe(401); + + const aged = await signAgedCallbackToken(WS_ACTIVE, 'workspace', 23 * HOUR); + const renewal = await renew(WS_ACTIVE, aged, { nodeId: NODE_ID, nodeToken }); + const { token } = (await renewal.json()) as { token: string }; + + const accepted = await SELF.fetch( + `https://api.test.example.com/api/workspaces/${WS_ACTIVE}/runtime`, + { headers: { Authorization: `Bearer ${token}` } } + ); + expect(accepted.status).toBe(200); + expect(await accepted.json()).toMatchObject({ workspaceId: WS_ACTIVE, nodeId: NODE_ID }); + }); +}); + +describe('mintWorkspaceCallbackTokenForNodeDelivery', () => { + it('mints a workspace token only for an active workspace bound to the target node', async () => { + const token = await mintWorkspaceCallbackTokenForNodeDelivery(testEnv, { + workspaceId: WS_ACTIVE, + nodeId: NODE_ID, + }); + expect(token).not.toBeNull(); + await expect( + verifyCallbackToken(token as string, testEnv, { expectedScope: 'workspace' }) + ).resolves.toEqual({ workspace: WS_ACTIVE, type: 'callback', scope: 'workspace' }); + }); + + it.each([ + ['workspace bound to another node', WS_ON_OTHER_NODE, NODE_ID], + ['deleted workspace', WS_DELETED, NODE_ID], + ['stopped workspace', WS_STOPPED, NODE_ID], + ['workspace on a stopped node', WS_ON_STOPPED_NODE, STOPPED_NODE_ID], + ['node owned by another user', WS_OTHER_OWNER_NODE, OTHER_OWNER_NODE_ID], + ['missing workspace', WS_MISSING, NODE_ID], + ])('delivers nothing for a %s', async (_label, workspaceId, nodeId) => { + await expect( + mintWorkspaceCallbackTokenForNodeDelivery(testEnv, { workspaceId, nodeId }) + ).resolves.toBeNull(); + }); +}); + +describe('hibernateAgentSessionOnNode workspace token delivery', () => { + const fetchMock = vi.fn(); + + afterEach(() => { + vi.unstubAllGlobals(); + fetchMock.mockReset(); + }); + + async function capturedHibernateBody(workspaceId: string, nodeId: string) { + fetchMock.mockResolvedValue( + new Response(JSON.stringify({ status: 'pending', accepted: true }), { status: 202 }) + ); + vi.stubGlobal('fetch', fetchMock); + + await hibernateAgentSessionOnNode(nodeId, workspaceId, 'agent-session-1', testEnv, USER_ID, { + chatSessionId: 'chat-1', + runtime: 'vm', + background: true, + }); + + expect(fetchMock).toHaveBeenCalledTimes(1); + const [url, init] = fetchMock.mock.calls[0] as [string, RequestInit]; + expect(String(url)).toContain(`/workspaces/${workspaceId}/agent-sessions/agent-session-1/hibernate`); + return JSON.parse(String(init.body)) as Record; + } + + it('carries a fresh workspace token to the node hosting an active workspace', async () => { + const body = await capturedHibernateBody(WS_ACTIVE, NODE_ID); + + expect(body).toMatchObject({ chatSessionId: 'chat-1', runtime: 'vm', background: true }); + expect(typeof body.workspaceCallbackToken).toBe('string'); + await expect( + verifyCallbackToken(body.workspaceCallbackToken as string, testEnv, { + expectedScope: 'workspace', + }) + ).resolves.toMatchObject({ workspace: WS_ACTIVE }); + }); + + it('sends the request without a token when the workspace is not bound to that node', async () => { + const body = await capturedHibernateBody(WS_ON_OTHER_NODE, NODE_ID); + + expect(body).toEqual({ chatSessionId: 'chat-1', runtime: 'vm', background: true }); + }); +}); From 4a90f5b01771d451b36b72a0a4a4e849057fb4dd Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 08:42:43 +0000 Subject: [PATCH 03/19] fix(api): deliver workspace tokens to VM nodes only and re-check incarnation Instant runtimes get a fresh token per cold wake; a wall-clock token pushed to a container generation would defeat the stale-callback guard, so hibernate delivery is VM-only. The delivery mint re-reads the workspace binding after signing and withholds the token if a delete or move won the race. Unit tests drive those races against real SQLite, plus legacy unscoped proofs, the expiry boundary and invalid refresh-ratio config. Co-Authored-By: Claude Opus 5.5 --- apps/api/src/routes/_stale-callback-guard.ts | 4 +- .../workspace-callback-token-renewal.ts | 78 ++++-- .../unit/routes/stale-callback-guard.test.ts | 14 +- .../workspace-callback-token-renewal.test.ts | 255 ++++++++++++++++++ .../workspace-callback-token-renewal.test.ts | 20 +- 5 files changed, 336 insertions(+), 35 deletions(-) create mode 100644 apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts diff --git a/apps/api/src/routes/_stale-callback-guard.ts b/apps/api/src/routes/_stale-callback-guard.ts index 1669b84105..7e1548c736 100644 --- a/apps/api/src/routes/_stale-callback-guard.ts +++ b/apps/api/src/routes/_stale-callback-guard.ts @@ -36,9 +36,7 @@ export const DEFAULT_INSTANT_STALE_CALLBACK_MARGIN_MS = 60_000; export function getInstantStaleCallbackMarginMs(env: Env): number { const raw = env.INSTANT_STALE_CALLBACK_MARGIN_MS; const parsed = raw ? Number.parseInt(raw, 10) : Number.NaN; - return Number.isFinite(parsed) && parsed >= 0 - ? parsed - : DEFAULT_INSTANT_STALE_CALLBACK_MARGIN_MS; + return Number.isFinite(parsed) && parsed >= 0 ? parsed : DEFAULT_INSTANT_STALE_CALLBACK_MARGIN_MS; } /** diff --git a/apps/api/src/services/workspace-callback-token-renewal.ts b/apps/api/src/services/workspace-callback-token-renewal.ts index 1ed8ed721e..915fd279de 100644 --- a/apps/api/src/services/workspace-callback-token-renewal.ts +++ b/apps/api/src/services/workspace-callback-token-renewal.ts @@ -14,8 +14,11 @@ * workspace token, a workspace token copied out of a devcontainer cannot renew * itself, and an expired token is never renewed. * 2. Control-plane delivery (API -> VM agent, `mintWorkspaceCallbackTokenForNodeDelivery`): - * requests the control plane already sends over the node-management channel carry a - * freshly minted token, exactly as workspace create and cf-container restore do. + * requests the control plane already sends over the node-management channel to a VM + * node carry a freshly minted token, exactly as workspace create does. Instant + * (cf-container) runtimes are excluded: their container DO mints a fresh token on every + * cold wake, and a wall-clock token pushed to a container generation would defeat the + * Instant stale-callback guard (`routes/_stale-callback-guard.ts`). * * Both paths renew only while D1 says the workspace is active and bound to the node, so * deleting, stopping or reassigning a workspace still ends its callback authority. @@ -26,6 +29,7 @@ import { AppError, errors } from '../middleware/error'; import { assertWorkspaceAcceptsCallback, assertWorkspaceCallbackIdentityCurrent, + sameWorkspaceCallbackIdentity, WORKSPACE_CALLBACK_ACTIVE_STATUSES, type WorkspaceCallbackIdentitySnapshot, } from '../routes/workspaces/_helpers'; @@ -46,11 +50,11 @@ export const NODE_CALLBACK_FORBIDDEN = 'NODE_CALLBACK_FORBIDDEN'; interface WorkspaceRenewalBinding extends WorkspaceCallbackIdentitySnapshot { nodeUserId: string | null; + nodeRuntime: string | null; } export type WorkspaceCallbackTokenRenewalResult = - | { renewed: true; token: string; expiresAt: string | null } - | { renewed: false }; + { renewed: true; token: string; expiresAt: string | null } | { renewed: false }; async function loadWorkspaceRenewalBinding( env: Env, @@ -64,7 +68,8 @@ async function loadWorkspaceRenewalBinding( w.status AS status, w.node_id AS nodeId, n.status AS nodeStatus, - n.user_id AS nodeUserId + n.user_id AS nodeUserId, + n.runtime AS nodeRuntime FROM workspaces w LEFT JOIN nodes n ON n.id = w.node_id WHERE w.id = ? @@ -186,36 +191,63 @@ export async function renewWorkspaceCallbackToken( }; } +/** + * Why a binding may not receive a delivered token, or null when it may. Delivery is + * VM-only (see the file header for why Instant runtimes are excluded). + */ +function deliverySkipReason( + binding: WorkspaceRenewalBinding | null, + nodeId: string +): string | null { + if (!binding) return 'workspace_missing'; + if (!workspaceBoundToNode(binding, nodeId)) return 'not_bound_to_node'; + if (binding.nodeRuntime === 'cf-container') return 'instant_runtime'; + if (!WORKSPACE_CALLBACK_ACTIVE_STATUSES.has(binding.status)) return 'workspace_inactive'; + if (!binding.nodeStatus || nodeStatusTerminatesCallbacks(binding.nodeStatus)) { + return 'node_inactive'; + } + return null; +} + +function sameRenewalBinding(current: WorkspaceRenewalBinding, expected: WorkspaceRenewalBinding) { + return ( + sameWorkspaceCallbackIdentity(current, expected) && + current.nodeUserId === expected.nodeUserId && + current.nodeRuntime === expected.nodeRuntime + ); +} + /** * Mint a fresh workspace callback token for a control-plane request that is about to be * delivered to `nodeId` over the node-management channel. Returns null, and the caller * sends its request without a token exactly as before, unless D1 binds the workspace to - * that node and the workspace and node are still active. + * that VM node and the workspace and node are still active, both before and after + * signing (rule 49), so a delete or move that wins the race gets no credential. */ export async function mintWorkspaceCallbackTokenForNodeDelivery( env: Env, input: { workspaceId: string; nodeId: string } ): Promise { + const skip = (reason: string) => { + log.info('workspace_callback_token.delivery_skipped', { + workspaceId: input.workspaceId, + nodeId: input.nodeId, + reason, + }); + return null; + }; try { const binding = await loadWorkspaceRenewalBinding(env, input.workspaceId); - const skipReason = !binding - ? 'workspace_missing' - : !workspaceBoundToNode(binding, input.nodeId) - ? 'not_bound_to_node' - : !WORKSPACE_CALLBACK_ACTIVE_STATUSES.has(binding.status) - ? 'workspace_inactive' - : !binding.nodeStatus || nodeStatusTerminatesCallbacks(binding.nodeStatus) - ? 'node_inactive' - : null; - if (skipReason) { - log.info('workspace_callback_token.delivery_skipped', { - workspaceId: input.workspaceId, - nodeId: input.nodeId, - reason: skipReason, - }); - return null; + const skipReason = deliverySkipReason(binding, input.nodeId); + if (skipReason || !binding) return skip(skipReason ?? 'workspace_missing'); + + const token = await signCallbackToken(input.workspaceId, env); + + const current = await loadWorkspaceRenewalBinding(env, input.workspaceId); + if (!current || !sameRenewalBinding(current, binding)) { + return skip('incarnation_changed'); } - return await signCallbackToken(input.workspaceId, env); + return token; } catch (err) { log.warn('workspace_callback_token.delivery_mint_failed', { workspaceId: input.workspaceId, diff --git a/apps/api/tests/unit/routes/stale-callback-guard.test.ts b/apps/api/tests/unit/routes/stale-callback-guard.test.ts index 97e1f83f1e..5194f2735c 100644 --- a/apps/api/tests/unit/routes/stale-callback-guard.test.ts +++ b/apps/api/tests/unit/routes/stale-callback-guard.test.ts @@ -35,9 +35,9 @@ describe('callbackTokenIssuedAtMs', () => { }); it('returns the preserved generation (gen_iat) of a renewed token, not its renewal iat', () => { - expect(callbackTokenIssuedAtMs(jwtWith({ iat: IAT_SECONDS + 86_400, gen_iat: IAT_SECONDS }))).toBe( - IAT_MS - ); + expect( + callbackTokenIssuedAtMs(jwtWith({ iat: IAT_SECONDS + 86_400, gen_iat: IAT_SECONDS })) + ).toBe(IAT_MS); }); it('falls back to iat when gen_iat is malformed', () => { @@ -76,12 +76,16 @@ describe('callbackTokenIssuedAtMs', () => { describe('getInstantStaleCallbackMarginMs', () => { it('defaults when unset', () => { - expect(getInstantStaleCallbackMarginMs({} as Env)).toBe(DEFAULT_INSTANT_STALE_CALLBACK_MARGIN_MS); + expect(getInstantStaleCallbackMarginMs({} as Env)).toBe( + DEFAULT_INSTANT_STALE_CALLBACK_MARGIN_MS + ); }); it('honours a valid override', () => { expect( - getInstantStaleCallbackMarginMs({ INSTANT_STALE_CALLBACK_MARGIN_MS: '30000' } as unknown as Env) + getInstantStaleCallbackMarginMs({ + INSTANT_STALE_CALLBACK_MARGIN_MS: '30000', + } as unknown as Env) ).toBe(30_000); }); diff --git a/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts b/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts new file mode 100644 index 0000000000..e7db3ad5af --- /dev/null +++ b/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts @@ -0,0 +1,255 @@ +/** + * Workspace callback token renewal and delivery against a real SQLite engine, with a D1 + * wrapper that can mutate the database between the reads a single call performs. This is + * how the race cases are driven: a delete or move that lands after the binding was read + * but before the credential leaves must suppress it (rule 49). + * + * Route-level behaviour (real Worker, real auth wiring) lives in + * tests/workers/workspace-callback-token-renewal.test.ts. + */ +import Database from 'better-sqlite3'; +import { decodeJwt, exportPKCS8, exportSPKI, generateKeyPair, importPKCS8, SignJWT } from 'jose'; +import { beforeAll, beforeEach, describe, expect, it } from 'vitest'; + +import * as schema from '../../../src/db/schema'; +import type { Env } from '../../../src/env'; +import { AppError } from '../../../src/middleware/error'; +import { CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM } from '../../../src/services/callback-token-claims'; +import { signNodeCallbackToken } from '../../../src/services/jwt'; +import { + mintWorkspaceCallbackTokenForNodeDelivery, + renewWorkspaceCallbackToken, +} from '../../../src/services/workspace-callback-token-renewal'; +import { createSchemaTables, createSqliteD1 } from '../../helpers/sqlite-d1'; + +const HOUR = 3600; +const DAY = 24 * HOUR; +const BASE_DOMAIN = 'example.com'; +const USER = 'user-1'; +const NODE = 'node-1'; +const OTHER_NODE = 'node-2'; +const WS = 'ws-1'; + +let privateKeyPem: string; +let publicKeyPem: string; + +beforeAll(async () => { + const { privateKey, publicKey } = await generateKeyPair('RS256', { extractable: true }); + privateKeyPem = await exportPKCS8(privateKey); + publicKeyPem = await exportSPKI(publicKey); +}); + +let sqlite: Database.Database; +let onPrepare: ((sql: string) => void) | null; + +function makeEnv(overrides: Partial> = {}): Env { + const d1 = createSqliteD1(sqlite); + const racing = new Proxy(d1, { + get(target, prop, receiver) { + if (prop === 'prepare') { + return (sql: string) => { + onPrepare?.(sql); + return target.prepare(sql); + }; + } + return Reflect.get(target, prop, receiver); + }, + }); + return { + DATABASE: racing, + BASE_DOMAIN, + JWT_PRIVATE_KEY: privateKeyPem, + JWT_PUBLIC_KEY: publicKeyPem, + ...overrides, + } as unknown as Env; +} + +/** Same claim set as production signers, issued `ageSeconds` ago. */ +async function agedToken( + subject: string, + scope: 'workspace' | 'node' | null, + ageSeconds: number, + lifetimeSeconds = DAY +): Promise { + const key = await importPKCS8(privateKeyPem, 'RS256'); + const issuedAt = Math.floor(Date.now() / 1000) - ageSeconds; + return new SignJWT({ workspace: subject, type: 'callback', ...(scope ? { scope } : {}) }) + .setProtectedHeader({ alg: 'RS256' }) + .setIssuer(`https://api.${BASE_DOMAIN}`) + .setSubject(subject) + .setAudience('workspace-callback') + .setIssuedAt(issuedAt) + .setExpirationTime(issuedAt + lifetimeSeconds) + .sign(key); +} + +function seed(): void { + const insertNode = sqlite.prepare( + 'INSERT INTO nodes (id, user_id, name, status, runtime) VALUES (?, ?, ?, ?, ?)' + ); + insertNode.run(NODE, USER, 'node-1', 'running', 'vm'); + insertNode.run(OTHER_NODE, USER, 'node-2', 'running', 'vm'); + sqlite + .prepare( + `INSERT INTO workspaces (id, user_id, node_id, project_id, chat_session_id, status, name) + VALUES (?, ?, ?, NULL, NULL, 'running', 'ws')` + ) + .run(WS, USER, NODE); +} + +function bindingReadCount(): { count: number; mutateOn: (n: number, sql: string) => void } { + const state = { count: 0, mutateOn: (_n: number, _sql: string) => undefined as void }; + onPrepare = (sql) => { + if (sql.includes('FROM workspaces w') && sql.includes('LEFT JOIN nodes n')) { + state.count += 1; + state.mutateOn(state.count, sql); + } + }; + return state; +} + +beforeEach(() => { + sqlite = new Database(':memory:'); + createSchemaTables(sqlite, [schema.workspaces, schema.nodes]); + seed(); + onPrepare = null; +}); + +describe('mintWorkspaceCallbackTokenForNodeDelivery', () => { + it('mints when the binding is unchanged across both reads', async () => { + const reads = bindingReadCount(); + const token = await mintWorkspaceCallbackTokenForNodeDelivery(makeEnv(), { + workspaceId: WS, + nodeId: NODE, + }); + expect(reads.count).toBe(2); + expect(token).not.toBeNull(); + expect(decodeJwt(token as string)).toMatchObject({ workspace: WS, scope: 'workspace' }); + }); + + it('withholds the token when the workspace is deleted after the first read', async () => { + const reads = bindingReadCount(); + reads.mutateOn = (n) => { + if (n === 2) sqlite.prepare("UPDATE workspaces SET status = 'deleted' WHERE id = ?").run(WS); + }; + await expect( + mintWorkspaceCallbackTokenForNodeDelivery(makeEnv(), { workspaceId: WS, nodeId: NODE }) + ).resolves.toBeNull(); + }); + + it('withholds the token when the workspace moves to another node after the first read', async () => { + const reads = bindingReadCount(); + reads.mutateOn = (n) => { + if (n === 2) + sqlite.prepare('UPDATE workspaces SET node_id = ? WHERE id = ?').run(OTHER_NODE, WS); + }; + await expect( + mintWorkspaceCallbackTokenForNodeDelivery(makeEnv(), { workspaceId: WS, nodeId: NODE }) + ).resolves.toBeNull(); + }); + + it('never delivers to an Instant (cf-container) runtime', async () => { + sqlite.prepare("UPDATE nodes SET runtime = 'cf-container' WHERE id = ?").run(NODE); + await expect( + mintWorkspaceCallbackTokenForNodeDelivery(makeEnv(), { workspaceId: WS, nodeId: NODE }) + ).resolves.toBeNull(); + }); + + it('returns null instead of throwing when the database read fails', async () => { + onPrepare = () => { + throw new Error('D1 unavailable'); + }; + await expect( + mintWorkspaceCallbackTokenForNodeDelivery(makeEnv(), { workspaceId: WS, nodeId: NODE }) + ).resolves.toBeNull(); + }); +}); + +describe('renewWorkspaceCallbackToken', () => { + async function renew( + workspaceToken: string, + env: Env = makeEnv(), + nodeToken?: string + ): Promise<{ renewed: boolean; token?: string }> { + return renewWorkspaceCallbackToken(env, { + workspaceId: WS, + workspaceToken, + nodeId: NODE, + nodeToken: nodeToken ?? (await signNodeCallbackToken(NODE, env)), + }); + } + + async function rejection(promise: Promise): Promise { + const error = await promise.then( + () => null, + (err: unknown) => err + ); + expect(error).toBeInstanceOf(AppError); + return error as AppError; + } + + it('suppresses the renewed credential when the workspace is deleted while minting', async () => { + const env = makeEnv(); + const nodeToken = await signNodeCallbackToken(NODE, env); + onPrepare = (sql) => { + // The identity re-read at the secret-delivery boundary is drizzle's select. + if (sql.includes('from "workspaces"') && sql.includes('left join "nodes"')) { + sqlite.prepare("UPDATE workspaces SET status = 'deleted' WHERE id = ?").run(WS); + } + }; + + const error = await rejection( + renew(await agedToken(WS, 'workspace', 13 * HOUR), env, nodeToken) + ); + expect(error.statusCode).toBe(410); + }); + + it('keeps the original generation of a legacy token without gen_iat', async () => { + const aged = await agedToken(WS, 'workspace', 13 * HOUR); + const result = await renew(aged); + expect(result.renewed).toBe(true); + expect(decodeJwt(result.token as string)[CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM]).toBe( + decodeJwt(aged).iat + ); + }); + + it('rejects legacy unscoped tokens as either proof', async () => { + const env = makeEnv(); + const legacyWorkspace = await agedToken(WS, null, 13 * HOUR); + const legacyNode = await agedToken(NODE, null, HOUR); + + const asWorkspaceProof = await rejection(renew(legacyWorkspace, env)); + expect([asWorkspaceProof.statusCode, asWorkspaceProof.error]).toEqual([403, 'FORBIDDEN']); + + const asNodeProof = await rejection( + renew(await agedToken(WS, 'workspace', 13 * HOUR), env, legacyNode) + ); + expect([asNodeProof.statusCode, asNodeProof.error]).toEqual([403, 'NODE_CALLBACK_FORBIDDEN']); + }); + + it('renews up to the expiry boundary and never past it', async () => { + const justValid = await agedToken(WS, 'workspace', DAY - 60); + expect((await renew(justValid)).renewed).toBe(true); + + const justExpired = await agedToken(WS, 'workspace', DAY + 1); + const error = await rejection(renew(justExpired)); + expect([error.statusCode, error.error]).toEqual([401, 'UNAUTHORIZED']); + }); + + it.each([ + ['unset (default 0.5)', undefined, 11, 13], + ['unparseable (falls back to 0.5)', 'not-a-number', 11, 13], + ['above the maximum (clamped to 0.9)', '0.99', 21, 22], + ['below the minimum (clamped to 0.1)', '0', 2, 3], + ])('applies the refresh ratio when %s', async (_label, ratio, notDueAgeHours, dueAgeHours) => { + const env = makeEnv( + ratio === undefined ? {} : { CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO: ratio } + ); + expect(await renew(await agedToken(WS, 'workspace', notDueAgeHours * HOUR), env)).toEqual({ + renewed: false, + }); + expect((await renew(await agedToken(WS, 'workspace', dueAgeHours * HOUR), env)).renewed).toBe( + true + ); + }); +}); diff --git a/apps/api/tests/workers/workspace-callback-token-renewal.test.ts b/apps/api/tests/workers/workspace-callback-token-renewal.test.ts index 2de68b252c..a320cd30df 100644 --- a/apps/api/tests/workers/workspace-callback-token-renewal.test.ts +++ b/apps/api/tests/workers/workspace-callback-token-renewal.test.ts @@ -38,6 +38,7 @@ const NODE_ID = `${PREFIX}-node`; const OTHER_NODE_ID = `${PREFIX}-node-2`; const OTHER_OWNER_NODE_ID = `${PREFIX}-node-other-owner`; const STOPPED_NODE_ID = `${PREFIX}-node-stopped`; +const INSTANT_NODE_ID = `${PREFIX}-node-instant`; const WS_ACTIVE = `${PREFIX}-ws-active`; const WS_CREATING = `${PREFIX}-ws-creating`; @@ -47,6 +48,7 @@ const WS_STOPPED = `${PREFIX}-ws-stopped`; const WS_ON_STOPPED_NODE = `${PREFIX}-ws-stopped-node`; const WS_OTHER_OWNER_NODE = `${PREFIX}-ws-other-owner-node`; const WS_MISSING = `${PREFIX}-ws-missing`; +const WS_INSTANT = `${PREFIX}-ws-instant`; let nodeToken: string; let otherNodeToken: string; @@ -122,6 +124,10 @@ beforeAll(async () => { await seedNode(OTHER_NODE_ID, USER_ID); await seedNode(OTHER_OWNER_NODE_ID, OTHER_USER_ID); await seedNode(STOPPED_NODE_ID, USER_ID, { status: 'stopped', healthStatus: 'unhealthy' }); + await seedNode(INSTANT_NODE_ID, USER_ID); + await testEnv.DATABASE.prepare("UPDATE nodes SET runtime = 'cf-container' WHERE id = ?") + .bind(INSTANT_NODE_ID) + .run(); await seedWorkspace(WS_ACTIVE, NODE_ID, USER_ID, { status: 'running' }); await seedWorkspace(WS_CREATING, NODE_ID, USER_ID, { status: 'creating' }); @@ -131,6 +137,7 @@ beforeAll(async () => { await seedWorkspace(WS_ON_STOPPED_NODE, STOPPED_NODE_ID, USER_ID, { status: 'running' }); // Corrupt binding: placement never puts a workspace on another user's node. await seedWorkspace(WS_OTHER_OWNER_NODE, OTHER_OWNER_NODE_ID, USER_ID, { status: 'running' }); + await seedWorkspace(WS_INSTANT, INSTANT_NODE_ID, USER_ID, { status: 'running' }); nodeToken = await signNodeCallbackToken(NODE_ID, testEnv); otherNodeToken = await signNodeCallbackToken(OTHER_NODE_ID, testEnv); @@ -174,9 +181,7 @@ describe('POST /api/workspaces/:id/callback-token/renew', () => { expect(response.status).toBe(200); const body = (await response.json()) as { token: string }; - expect(decodeJwt(body.token)[CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM]).toBe( - originalIssuedAt - ); + expect(decodeJwt(body.token)[CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM]).toBe(originalIssuedAt); }); it('does not mint while the token is younger than the refresh threshold', async () => { @@ -372,6 +377,11 @@ describe('mintWorkspaceCallbackTokenForNodeDelivery', () => { ['workspace on a stopped node', WS_ON_STOPPED_NODE, STOPPED_NODE_ID], ['node owned by another user', WS_OTHER_OWNER_NODE, OTHER_OWNER_NODE_ID], ['missing workspace', WS_MISSING, NODE_ID], + [ + 'Instant (cf-container) runtime, which gets a fresh token per cold wake', + WS_INSTANT, + INSTANT_NODE_ID, + ], ])('delivers nothing for a %s', async (_label, workspaceId, nodeId) => { await expect( mintWorkspaceCallbackTokenForNodeDelivery(testEnv, { workspaceId, nodeId }) @@ -401,7 +411,9 @@ describe('hibernateAgentSessionOnNode workspace token delivery', () => { expect(fetchMock).toHaveBeenCalledTimes(1); const [url, init] = fetchMock.mock.calls[0] as [string, RequestInit]; - expect(String(url)).toContain(`/workspaces/${workspaceId}/agent-sessions/agent-session-1/hibernate`); + expect(String(url)).toContain( + `/workspaces/${workspaceId}/agent-sessions/agent-session-1/hibernate` + ); return JSON.parse(String(init.body)) as Record; } From aec1391c2648c53c484ef3c4f7a3566c1503b695 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 08:53:09 +0000 Subject: [PATCH 04/19] feat(vm-agent): renew workspace callback tokens and propagate every change - After each successful heartbeat, renew workspace tokens past WORKSPACE_CALLBACK_TOKEN_REFRESH_RATIO of their lifetime with the dual-proof renewal route; latch refusals per token, back off transient failures, and install the result by compare-and-swap. - Persist every adopted token (renewal or control-plane delivery) before publishing it, never adopt a delivered token that expires earlier than the current one, and propagate changes to the message reporter and SessionHosts. - SessionHost reads its callback token through a lock-free accessor. - A 401 on message persistence no longer deletes the outbox: the reporter holds rows, resends at once after a rotation, resumes on a replacement token, and reports a pause longer than MSG_AUTH_RENEWAL_WAIT on the node error channel. Co-Authored-By: Claude Opus 5.5 --- .../vm-agent/internal/acp/session_host.go | 4 + .../acp/session_host_callback_token.go | 28 ++ .../internal/acp/session_host_form.go | 2 +- .../acp/session_host_interaction_transport.go | 6 +- .../internal/acp/session_host_interactions.go | 2 +- .../internal/acp/session_host_reporting.go | 6 +- .../internal/acp/session_host_startup.go | 12 +- .../vm-agent/internal/acp/session_host_url.go | 4 +- .../internal/acp/session_host_usage.go | 2 +- .../internal/config/callback_token_renewal.go | 54 +++ packages/vm-agent/internal/config/config.go | 6 + .../vm-agent/internal/config/config_load.go | 5 + .../vm-agent/internal/messagereport/config.go | 18 + .../internal/messagereport/credential.go | 145 +++++++ .../internal/messagereport/credential_test.go | 350 +++++++++++++++ .../internal/messagereport/reporter.go | 29 +- .../internal/messagereport/reporter_test.go | 7 +- .../vm-agent/internal/messagereport/sender.go | 19 +- packages/vm-agent/internal/server/health.go | 2 + packages/vm-agent/internal/server/server.go | 2 + .../workspace_callback_token_renewal.go | 407 ++++++++++++++++++ .../internal/server/workspace_routing.go | 7 +- 22 files changed, 1092 insertions(+), 25 deletions(-) create mode 100644 packages/vm-agent/internal/acp/session_host_callback_token.go create mode 100644 packages/vm-agent/internal/config/callback_token_renewal.go create mode 100644 packages/vm-agent/internal/messagereport/credential.go create mode 100644 packages/vm-agent/internal/messagereport/credential_test.go create mode 100644 packages/vm-agent/internal/server/workspace_callback_token_renewal.go diff --git a/packages/vm-agent/internal/acp/session_host.go b/packages/vm-agent/internal/acp/session_host.go index a9bc469880..bc1c5e8562 100644 --- a/packages/vm-agent/internal/acp/session_host.go +++ b/packages/vm-agent/internal/acp/session_host.go @@ -216,6 +216,10 @@ type SessionHost struct { // credentialAttribution stores non-secret server-selected credential identity // for usage callbacks. It is lock-free so SessionUpdate never waits on h.mu. credentialAttribution atomic.Value + // renewedCallbackToken holds a workspace callback token delivered after the + // host was created (SetCallbackToken). Lock-free like the fields above: + // control-plane reporting runs on the ACP notification goroutine. + renewedCallbackToken atomic.Value // string // Credential injection metadata (set during startAgent, read during stop). // These track whether the agent used file-based credential injection so diff --git a/packages/vm-agent/internal/acp/session_host_callback_token.go b/packages/vm-agent/internal/acp/session_host_callback_token.go new file mode 100644 index 0000000000..817c6ef506 --- /dev/null +++ b/packages/vm-agent/internal/acp/session_host_callback_token.go @@ -0,0 +1,28 @@ +package acp + +import "strings" + +// callbackToken returns the workspace callback token for control-plane calls: +// the latest one delivered through SetCallbackToken, else the token the host was +// created with. Every control-plane request and every agent process start reads +// it here, so a renewal reaches activity, usage, interaction, agent-key and +// runtime-asset calls, and the credentials injected into the next agent start. +// +// Lock-free: callers include the ACP notification goroutine, which must never +// block on h.mu (see the mirror fields on SessionHost and .claude/rules/46). +func (h *SessionHost) callbackToken() string { + if token, ok := h.renewedCallbackToken.Load().(string); ok && token != "" { + return token + } + return h.config.CallbackToken +} + +// SetCallbackToken replaces the workspace callback token used by this host's +// later control-plane calls. Empty tokens are ignored. An agent process that is +// already running keeps the credentials it was started with (for example the +// platform AI proxy key); it picks up the new token on its next (re)start. +func (h *SessionHost) SetCallbackToken(token string) { + if token = strings.TrimSpace(token); token != "" { + h.renewedCallbackToken.Store(token) + } +} diff --git a/packages/vm-agent/internal/acp/session_host_form.go b/packages/vm-agent/internal/acp/session_host_form.go index 5d82e46210..4c84f5938e 100644 --- a/packages/vm-agent/internal/acp/session_host_form.go +++ b/packages/vm-agent/internal/acp/session_host_form.go @@ -64,7 +64,7 @@ func (h *SessionHost) requestForm(ctx context.Context, generation string, !config.Enabled || !config.FormsEnabled || config.validate() != nil || generation == "" || h.config.ProjectID == "" || h.config.WorkspaceID == "" || h.config.SessionID == "" || h.config.RuntimeIdentity == "" || - h.config.CallbackToken == "" || h.config.ControlPlaneURL == "" { + h.callbackToken() == "" || h.config.ControlPlaneURL == "" { slog.Info("acp_interaction.form_cancelled", "reason", "unsupported") return cancelledFormResponse(), nil } diff --git a/packages/vm-agent/internal/acp/session_host_interaction_transport.go b/packages/vm-agent/internal/acp/session_host_interaction_transport.go index 25afef098d..e2bc099603 100644 --- a/packages/vm-agent/internal/acp/session_host_interaction_transport.go +++ b/packages/vm-agent/internal/acp/session_host_interaction_transport.go @@ -34,7 +34,7 @@ func (h *SessionHost) createAcpInteraction(ctx context.Context, request acpInter if err != nil { return acpInteractionCreateRejected, fmt.Errorf("build ACP interaction create: %w", err) } - httpRequest.Header.Set("Authorization", "Bearer "+h.config.CallbackToken) + httpRequest.Header.Set("Authorization", "Bearer "+h.callbackToken()) httpRequest.Header.Set("Content-Type", "application/json") response, err := h.httpClient().Do(httpRequest) if err != nil { @@ -59,7 +59,7 @@ func (h *SessionHost) createAcpInteraction(ctx context.Context, request acpInter func (h *SessionHost) settleAcpInteraction(request acpInteractionSettleRequest, deadline time.Time) { config := h.acpInteractionConfigSnapshot() - if !config.Enabled || h.config.CallbackToken == "" || h.config.ControlPlaneURL == "" { + if !config.Enabled || h.callbackToken() == "" || h.config.ControlPlaneURL == "" { return } maxDuration := time.Duration(config.MaxDeadlineMs) * time.Millisecond @@ -99,7 +99,7 @@ func (h *SessionHost) settleAcpInteraction(request acpInteractionSettleRequest, if buildErr != nil { return } - httpRequest.Header.Set("Authorization", "Bearer "+h.config.CallbackToken) + httpRequest.Header.Set("Authorization", "Bearer "+h.callbackToken()) httpRequest.Header.Set("Content-Type", "application/json") response, sendErr := h.httpClient().Do(httpRequest) if sendErr == nil { diff --git a/packages/vm-agent/internal/acp/session_host_interactions.go b/packages/vm-agent/internal/acp/session_host_interactions.go index abd017b0ae..d9c60a8c5d 100644 --- a/packages/vm-agent/internal/acp/session_host_interactions.go +++ b/packages/vm-agent/internal/acp/session_host_interactions.go @@ -510,7 +510,7 @@ func (h *SessionHost) requestPermission( config := h.acpInteractionConfigSnapshot() if !config.Enabled || config.validate() != nil || generation == "" || h.config.ProjectID == "" || h.config.WorkspaceID == "" || h.config.SessionID == "" || - h.config.RuntimeIdentity == "" || h.config.CallbackToken == "" || h.config.ControlPlaneURL == "" { + h.config.RuntimeIdentity == "" || h.callbackToken() == "" || h.config.ControlPlaneURL == "" { slog.Info("acp_interaction.permission_cancelled", "reason", "unsupported") return cancelledPermissionResponse(), nil } diff --git a/packages/vm-agent/internal/acp/session_host_reporting.go b/packages/vm-agent/internal/acp/session_host_reporting.go index 4ead373219..468e3466a6 100644 --- a/packages/vm-agent/internal/acp/session_host_reporting.go +++ b/packages/vm-agent/internal/acp/session_host_reporting.go @@ -78,7 +78,7 @@ func (h *SessionHost) fetchAgentKey(ctx context.Context, agentType string) (*age return nil, fmt.Errorf("failed to create request: %w", err) } req.Header.Set("Content-Type", "application/json") - req.Header.Set("Authorization", "Bearer "+h.config.CallbackToken) + req.Header.Set("Authorization", "Bearer "+h.callbackToken()) resp, err := h.httpClient().Do(req) if err != nil { @@ -152,7 +152,7 @@ func (h *SessionHost) fetchAgentSettings(ctx context.Context, agentType string) return nil } req.Header.Set("Content-Type", "application/json") - req.Header.Set("Authorization", "Bearer "+h.config.CallbackToken) + req.Header.Set("Authorization", "Bearer "+h.callbackToken()) resp, err := h.httpClient().Do(req) if err != nil { @@ -282,7 +282,7 @@ func (h *SessionHost) prepareActivityReport(activity string) (activityReportRequ projectID := h.config.ProjectID nodeID := h.config.NodeID controlPlaneURL := h.config.ControlPlaneURL - callbackToken := h.config.CallbackToken + callbackToken := h.callbackToken() sessionID := h.config.SessionID if projectID == "" || nodeID == "" || controlPlaneURL == "" || sessionID == "" { diff --git a/packages/vm-agent/internal/acp/session_host_startup.go b/packages/vm-agent/internal/acp/session_host_startup.go index e4dd7e5ef4..af867c3195 100644 --- a/packages/vm-agent/internal/acp/session_host_startup.go +++ b/packages/vm-agent/internal/acp/session_host_startup.go @@ -301,7 +301,7 @@ func (h *SessionHost) injectAuthFileCredential( func (h *SessionHost) codexRefreshProxyEnv(agentType string, cred *agentCredential) (string, bool) { if agentType != "openai-codex" || cred.credentialKind != "oauth-token" || - h.config.ControlPlaneURL == "" || h.config.CallbackToken == "" { + h.config.ControlPlaneURL == "" || h.callbackToken() == "" { return "", false } u, err := url.Parse(strings.TrimSuffix(h.config.ControlPlaneURL, "/") + "/api/auth/codex-refresh") @@ -311,7 +311,7 @@ func (h *SessionHost) codexRefreshProxyEnv(agentType string, cred *agentCredenti return "", false } q := url.Values{} - q.Set("token", h.config.CallbackToken) + q.Set("token", h.callbackToken()) u.RawQuery = q.Encode() return "CODEX_REFRESH_TOKEN_URL_OVERRIDE=" + u.String(), true } @@ -362,7 +362,7 @@ func (h *SessionHost) injectPlatformProxyCredential( settings *agentSettingsPayload, envVars []string, ) ([]string, *agentSettingsPayload, error) { - return h.injectProxyCredential(agentType, cred, settings, envVars, "platform AI proxy", h.config.CallbackToken, "callbackTokenLen") + return h.injectProxyCredential(agentType, cred, settings, envVars, "platform AI proxy", h.callbackToken(), "callbackTokenLen") } // injectProxyCredential is the shared implementation behind the passthrough and @@ -378,7 +378,7 @@ func (h *SessionHost) injectProxyCredential( credential string, credLenKey string, ) ([]string, *agentSettingsPayload, error) { - if h.config.CallbackToken == "" { + if h.callbackToken() == "" { return envVars, settings, fmt.Errorf("%s configured but CallbackToken is empty for workspace %s", label, h.config.WorkspaceID) } @@ -399,7 +399,7 @@ func (h *SessionHost) proxyBaseURL(cred *agentCredential) string { if cred == nil || cred.inferenceConfig == nil { return "" } - return strings.ReplaceAll(cred.inferenceConfig.BaseURL, "{wstoken}", h.config.CallbackToken) + return strings.ReplaceAll(cred.inferenceConfig.BaseURL, "{wstoken}", h.callbackToken()) } type proxyEnvDescriptor struct { @@ -501,7 +501,7 @@ func (h *SessionHost) writeCodexStartupConfig(ctx context.Context, cred *agentCr if err != nil { return fmt.Errorf("cannot start Codex: %w", err) } - proxyConfig := codexProxyProviderConfigFromCredential(cred, h.config.CallbackToken) + proxyConfig := codexProxyProviderConfigFromCredential(cred, h.callbackToken()) effort := "" if startup.settings != nil { effort = startup.settings.Effort diff --git a/packages/vm-agent/internal/acp/session_host_url.go b/packages/vm-agent/internal/acp/session_host_url.go index 22bd6c5c22..52bd8e6286 100644 --- a/packages/vm-agent/internal/acp/session_host_url.go +++ b/packages/vm-agent/internal/acp/session_host_url.go @@ -173,7 +173,7 @@ func (h *SessionHost) requestURL(ctx context.Context, generation string, if params.Url == nil || params.Form != nil || len(params.Url.Meta) != 0 || !config.Enabled || !config.URLsEnabled || config.validate() != nil || h.config.ProjectID == "" || h.config.WorkspaceID == "" || h.config.SessionID == "" || - h.config.RuntimeIdentity == "" || h.config.CallbackToken == "" || h.config.ControlPlaneURL == "" || + h.config.RuntimeIdentity == "" || h.callbackToken() == "" || h.config.ControlPlaneURL == "" || len(params.Url.ElicitationId) == 0 || len(utf16.Encode([]rune(params.Url.ElicitationId))) > config.URLElicitationIDMaxChars || len(params.Url.Message) > config.RequestMaxBytes { return acpsdk.NewUnstableCreateElicitationResponseCancel(), nil @@ -347,7 +347,7 @@ func (h *SessionHost) completeURLInteraction(entry acpUrlElicitation, elicitatio if err != nil { return } - req.Header.Set("Authorization", "Bearer "+h.config.CallbackToken) + req.Header.Set("Authorization", "Bearer "+h.callbackToken()) req.Header.Set("Content-Type", "application/json") resp, err := h.httpClient().Do(req) if err != nil { diff --git a/packages/vm-agent/internal/acp/session_host_usage.go b/packages/vm-agent/internal/acp/session_host_usage.go index ec43363777..771edf1706 100644 --- a/packages/vm-agent/internal/acp/session_host_usage.go +++ b/packages/vm-agent/internal/acp/session_host_usage.go @@ -109,7 +109,7 @@ func (h *SessionHost) prepareUsageReportWithAttribution(params acpsdk.SessionNot projectID := h.config.ProjectID nodeID := h.config.NodeID controlPlaneURL := h.config.ControlPlaneURL - callbackToken := h.config.CallbackToken + callbackToken := h.callbackToken() sessionID := h.config.SessionID if projectID == "" || nodeID == "" || controlPlaneURL == "" || sessionID == "" || callbackToken == "" { return usageReportRequest{}, false diff --git a/packages/vm-agent/internal/config/callback_token_renewal.go b/packages/vm-agent/internal/config/callback_token_renewal.go new file mode 100644 index 0000000000..a6d6f46544 --- /dev/null +++ b/packages/vm-agent/internal/config/callback_token_renewal.go @@ -0,0 +1,54 @@ +package config + +import "time" + +// Workspace callback token renewal defaults. Workspace-scoped callback tokens are +// minted by the control plane with a fixed lifetime (CALLBACK_TOKEN_EXPIRY_MS, +// default 24h); the agent renews each one through +// POST /api/workspaces/:id/callback-token/renew once it has used up +// WorkspaceCallbackTokenRefreshRatio of that lifetime. +const ( + // DefaultWorkspaceCallbackTokenRefreshRatio renews halfway through a token's + // lifetime, matching the control plane's CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO + // default, so a renewal outage has the second half of the lifetime to recover. + // Override via WORKSPACE_CALLBACK_TOKEN_REFRESH_RATIO. + DefaultWorkspaceCallbackTokenRefreshRatio = 0.5 + // MinWorkspaceCallbackTokenRefreshRatio and MaxWorkspaceCallbackTokenRefreshRatio + // bound the ratio the same way the control plane clamps its own. + MinWorkspaceCallbackTokenRefreshRatio = 0.1 + MaxWorkspaceCallbackTokenRefreshRatio = 0.9 + + // DefaultWorkspaceCallbackTokenRenewalTimeout bounds one renewal request. + // Override via WORKSPACE_CALLBACK_TOKEN_RENEWAL_TIMEOUT. + DefaultWorkspaceCallbackTokenRenewalTimeout = 15 * time.Second + // DefaultWorkspaceCallbackTokenRenewalRetryInitial is the first backoff after a + // transient renewal failure. Override via WORKSPACE_CALLBACK_TOKEN_RENEWAL_RETRY_INITIAL. + DefaultWorkspaceCallbackTokenRenewalRetryInitial = time.Minute + // DefaultWorkspaceCallbackTokenRenewalRetryMax caps the backoff, and is also the + // wait before asking again when the control plane says a token is not yet due. + // Override via WORKSPACE_CALLBACK_TOKEN_RENEWAL_RETRY_MAX. + DefaultWorkspaceCallbackTokenRenewalRetryMax = 30 * time.Minute +) + +// clampWorkspaceCallbackTokenRefreshRatio keeps a configured ratio inside the +// supported range; a non-positive or unparseable value falls back to the default. +func clampWorkspaceCallbackTokenRefreshRatio(ratio float64) float64 { + if ratio <= 0 || ratio != ratio { // ratio != ratio rejects NaN + return DefaultWorkspaceCallbackTokenRefreshRatio + } + if ratio < MinWorkspaceCallbackTokenRefreshRatio { + return MinWorkspaceCallbackTokenRefreshRatio + } + if ratio > MaxWorkspaceCallbackTokenRefreshRatio { + return MaxWorkspaceCallbackTokenRefreshRatio + } + return ratio +} + +// positiveDurationOr returns value when positive, else fallback. +func positiveDurationOr(value, fallback time.Duration) time.Duration { + if value > 0 { + return value + } + return fallback +} diff --git a/packages/vm-agent/internal/config/config.go b/packages/vm-agent/internal/config/config.go index b6d036fe12..7de24801d0 100644 --- a/packages/vm-agent/internal/config/config.go +++ b/packages/vm-agent/internal/config/config.go @@ -402,6 +402,12 @@ type Config struct { // Callback retry settings - configurable per constitution principle XI WorkspaceReadyCallbackTimeout time.Duration // HTTP timeout for workspace-ready retry callbacks (env: WORKSPACE_READY_CALLBACK_TIMEOUT, default: 30s) + // Workspace callback token renewal (see callback_token_renewal.go for defaults and env vars) + WorkspaceCallbackTokenRefreshRatio float64 + WorkspaceCallbackTokenRenewalTimeout time.Duration + WorkspaceCallbackTokenRenewalRetryInitial time.Duration + WorkspaceCallbackTokenRenewalRetryMax time.Duration + // Error reporting settings - configurable per constitution principle XI ErrorReportFlushInterval time.Duration // Background flush interval (default: 30s) ErrorReportMaxBatchSize int // Immediate flush threshold (default: 10) diff --git a/packages/vm-agent/internal/config/config_load.go b/packages/vm-agent/internal/config/config_load.go index 2a7a61dc1a..e3589620f5 100644 --- a/packages/vm-agent/internal/config/config_load.go +++ b/packages/vm-agent/internal/config/config_load.go @@ -279,6 +279,11 @@ func Load() (*Config, error) { // Callback retry settings - configurable per constitution principle XI WorkspaceReadyCallbackTimeout: getEnvDuration("WORKSPACE_READY_CALLBACK_TIMEOUT", DefaultWorkspaceReadyCallbackTimeout), + WorkspaceCallbackTokenRefreshRatio: clampWorkspaceCallbackTokenRefreshRatio(getEnvFloat("WORKSPACE_CALLBACK_TOKEN_REFRESH_RATIO", DefaultWorkspaceCallbackTokenRefreshRatio)), + WorkspaceCallbackTokenRenewalTimeout: positiveDurationOr(getEnvDuration("WORKSPACE_CALLBACK_TOKEN_RENEWAL_TIMEOUT", DefaultWorkspaceCallbackTokenRenewalTimeout), DefaultWorkspaceCallbackTokenRenewalTimeout), + WorkspaceCallbackTokenRenewalRetryInitial: positiveDurationOr(getEnvDuration("WORKSPACE_CALLBACK_TOKEN_RENEWAL_RETRY_INITIAL", DefaultWorkspaceCallbackTokenRenewalRetryInitial), DefaultWorkspaceCallbackTokenRenewalRetryInitial), + WorkspaceCallbackTokenRenewalRetryMax: positiveDurationOr(getEnvDuration("WORKSPACE_CALLBACK_TOKEN_RENEWAL_RETRY_MAX", DefaultWorkspaceCallbackTokenRenewalRetryMax), DefaultWorkspaceCallbackTokenRenewalRetryMax), + // Error reporting settings - configurable per constitution principle XI ErrorReportFlushInterval: getEnvDuration("ERROR_REPORT_FLUSH_INTERVAL", 30*time.Second), ErrorReportMaxBatchSize: getEnvInt("ERROR_REPORT_MAX_BATCH_SIZE", 10), diff --git a/packages/vm-agent/internal/messagereport/config.go b/packages/vm-agent/internal/messagereport/config.go index 9041992e87..00406d6f5b 100644 --- a/packages/vm-agent/internal/messagereport/config.go +++ b/packages/vm-agent/internal/messagereport/config.go @@ -13,6 +13,12 @@ import ( const DefaultResponseMaxBytes = 2048 +// DefaultAuthRenewalWait is the default for Config.AuthRenewalWait. A healthy agent +// renews its workspace token halfway through the token's lifetime, so a rejected +// token that no renewal or control-plane delivery replaces within this window is +// worth surfacing. Override via MSG_AUTH_RENEWAL_WAIT. +const DefaultAuthRenewalWait = 15 * time.Minute + // Config holds tunable parameters for the message reporter. // All values have sensible defaults; override via MSG_* environment variables. type Config struct { @@ -50,6 +56,16 @@ type Config struct { // diagnostics when the control plane rejects a batch. ResponseMaxBytes int + // AuthRenewalWait is how long delivery may stay paused on a rejected (401) + // callback token before the pause is surfaced through OnAuthRenewalWaitExceeded. + // Queued rows are kept either way; see credential.go. + AuthRenewalWait time.Duration + + // OnAuthRenewalWaitExceeded, when set, is called once per pause that outlasts + // AuthRenewalWait, so the owner can report it on a channel that does not depend + // on the rejected workspace token. + OnAuthRenewalWaitExceeded func(AuthRenewalWaitExceeded) + // Endpoint is the control plane URL (without trailing slash). // The batch endpoint will be: {Endpoint}/api/workspaces/{workspaceId}/messages Endpoint string @@ -80,6 +96,7 @@ func DefaultConfig() Config { RetryMaxElapsed: 5 * time.Minute, HTTPTimeout: 10 * time.Second, ResponseMaxBytes: DefaultResponseMaxBytes, + AuthRenewalWait: DefaultAuthRenewalWait, } } @@ -98,6 +115,7 @@ func LoadConfigFromEnv() Config { cfg.RetryMaxElapsed = envDuration("MSG_RETRY_MAX_ELAPSED", cfg.RetryMaxElapsed) cfg.HTTPTimeout = envDuration("MSG_HTTP_TIMEOUT", cfg.HTTPTimeout) cfg.ResponseMaxBytes = envInt("MSG_RESPONSE_MAX_BYTES", cfg.ResponseMaxBytes) + cfg.AuthRenewalWait = envDuration("MSG_AUTH_RENEWAL_WAIT", cfg.AuthRenewalWait) cfg.Endpoint = os.Getenv("CONTROL_PLANE_URL") cfg.WorkspaceID = os.Getenv("WORKSPACE_ID") diff --git a/packages/vm-agent/internal/messagereport/credential.go b/packages/vm-agent/internal/messagereport/credential.go new file mode 100644 index 0000000000..a6c02b1fcc --- /dev/null +++ b/packages/vm-agent/internal/messagereport/credential.go @@ -0,0 +1,145 @@ +package messagereport + +import ( + "errors" + "log/slog" + "time" +) + +// A 401 from the messages endpoint means the control plane rejected the +// workspace callback token, not that the session or workspace is gone (those are +// 204/403/404/410). The token can still be replaced: the VM agent renews it +// before it expires, and the control plane re-delivers one on hibernate. So a +// 401 never deletes the outbox. Instead the reporter: +// +// - resends at once when a renewal replaced the token while the request was in +// flight (the 401 was for a token that is no longer current); +// - otherwise sends nothing while the rejected token is still current, keeps +// the queued rows (bounded by OutboxMaxSize), and resumes as soon as SetToken +// installs a different token; +// - surfaces a pause that outlasts AuthRenewalWait once, through an error log +// and OnAuthRenewalWaitExceeded. +// +// The control plane authenticates before it reads the body and dedupes messages +// by id, so a held row that is resent can never be persisted twice. +// +// Held rows exist only in this VM's outbox. They are NOT durable transcript until +// delivered: a teardown while delivery is paused loses them. + +var ( + errAwaitingCredential = errors.New("messagereport: holding messages until the workspace callback token is replaced") + errCredentialRotated = errors.New("messagereport: workspace callback token was replaced during the request") +) + +// maxRotatedTokenRetriesPerFlush bounds immediate resends after a rotation, so +// a flush cannot spin if tokens keep changing under it. +const maxRotatedTokenRetriesPerFlush = 3 + +// credentialWait is the paused-delivery state. Guarded by Reporter.mu. +type credentialWait struct { + rejectedToken string + since time.Time + reported bool +} + +// resumeWith ends the pause when token differs from the rejected one. +func (w *credentialWait) resumeWith(token string) bool { + if w.rejectedToken == "" || token == "" || token == w.rejectedToken { + return false + } + *w = credentialWait{} + return true +} + +// AuthRenewalWaitExceeded describes a delivery pause that outlasted +// Config.AuthRenewalWait. +type AuthRenewalWaitExceeded struct { + WorkspaceID string + SessionID string + HeldMessages int + PausedFor time.Duration +} + +// credentialRejectedError is a 401 seen while sending one row on its own. +type credentialRejectedError struct { + responseBody string +} + +func (e credentialRejectedError) Error() string { + return "workspace callback token rejected: " + e.responseBody +} + +// credentialRejected handles a 401 for token. It returns errCredentialRotated +// when a renewal already replaced token (resend now), else pauses delivery and +// returns errAwaitingCredential. +func (r *Reporter) credentialRejected(token, responseBody string) error { + r.mu.Lock() + if r.authToken != token { + r.mu.Unlock() + return errCredentialRotated + } + first := r.credentialWait.rejectedToken != token + if first { + r.credentialWait = credentialWait{rejectedToken: token, since: r.now()} + } + wsID, sessionID := r.workspaceID, r.sessionID + r.mu.Unlock() + if first { + slog.Warn("messagereport: control plane rejected the workspace callback token; holding messages until it is replaced", + "workspaceId", wsID, + "sessionId", sessionID, + "responseBody", responseBody, + ) + } + return errAwaitingCredential +} + +// awaitingCredential reports whether delivery is paused on token, and surfaces +// the pause once when it has lasted AuthRenewalWait. +func (r *Reporter) awaitingCredential(token string) bool { + r.mu.Lock() + if r.credentialWait.rejectedToken == "" || r.credentialWait.rejectedToken != token { + r.mu.Unlock() + return false + } + pausedFor := r.now().Sub(r.credentialWait.since) + report := !r.credentialWait.reported && pausedFor >= r.cfg.AuthRenewalWait + if report { + r.credentialWait.reported = true + } + wsID, sessionID := r.workspaceID, r.sessionID + r.mu.Unlock() + if report { + r.reportAuthRenewalWaitExceeded(AuthRenewalWaitExceeded{ + WorkspaceID: wsID, + SessionID: sessionID, + HeldMessages: r.heldMessageCount(), + PausedFor: pausedFor, + }) + } + return true +} + +func (r *Reporter) reportAuthRenewalWaitExceeded(info AuthRenewalWaitExceeded) { + slog.Error("messagereport: message persistence paused; the rejected workspace callback token has not been replaced", + "workspaceId", info.WorkspaceID, + "sessionId", info.SessionID, + "heldMessages", info.HeldMessages, + "pausedFor", info.PausedFor.String(), + ) + if r.cfg.OnAuthRenewalWaitExceeded != nil { + r.cfg.OnAuthRenewalWaitExceeded(info) + } +} + +// heldMessageCount is a bounded count of queued rows, or -1 if it cannot be read. +func (r *Reporter) heldMessageCount() int { + var count int + if err := r.db.QueryRow( + "SELECT COUNT(*) FROM (SELECT 1 FROM message_outbox LIMIT ?)", + r.cfg.OutboxMaxSize+1, + ).Scan(&count); err != nil { + return -1 + } + return count +} diff --git a/packages/vm-agent/internal/messagereport/credential_test.go b/packages/vm-agent/internal/messagereport/credential_test.go new file mode 100644 index 0000000000..5162642570 --- /dev/null +++ b/packages/vm-agent/internal/messagereport/credential_test.go @@ -0,0 +1,350 @@ +package messagereport + +import ( + "encoding/json" + "net/http" + "net/http/httptest" + "strings" + "sync" + "testing" + "time" +) + +// tokenGatedControlPlane mirrors the real messages endpoint contract: it rejects +// an unknown token with 401 BEFORE reading the body (nothing is persisted), and +// otherwise persists each message id once, counting repeats as duplicates +// (project-data/messages.ts dedupes by id). +type tokenGatedControlPlane struct { + t *testing.T + mu sync.Mutex + valid map[string]bool + persisted map[string]int + duplicates int + tokens []string + batchLimit int + onRequest func(token string) +} + +func newTokenGatedControlPlane(t *testing.T, valid ...string) (*tokenGatedControlPlane, *httptest.Server) { + cp := &tokenGatedControlPlane{t: t, valid: map[string]bool{}, persisted: map[string]int{}} + for _, token := range valid { + cp.valid[token] = true + } + server := httptest.NewServer(http.HandlerFunc(cp.serve)) + t.Cleanup(server.Close) + return cp, server +} + +func (cp *tokenGatedControlPlane) serve(w http.ResponseWriter, r *http.Request) { + token := strings.TrimPrefix(r.Header.Get("Authorization"), "Bearer ") + cp.mu.Lock() + hook := cp.onRequest + cp.mu.Unlock() + if hook != nil { + hook(token) + } + + cp.mu.Lock() + cp.tokens = append(cp.tokens, token) + valid := cp.valid[token] + cp.mu.Unlock() + if !valid { + w.WriteHeader(http.StatusUnauthorized) + _, _ = w.Write([]byte(`{"error":"UNAUTHORIZED","message":"Invalid or expired callback token"}`)) + return + } + + ids := requestMessageIDs(cp.t, r) + if cp.batchLimit > 0 && len(ids) > cp.batchLimit { + writePayloadTooLarge(w) + return + } + cp.mu.Lock() + persisted, duplicates := 0, 0 + for _, id := range ids { + if cp.persisted[id] > 0 { + duplicates++ + continue + } + cp.persisted[id] = 1 + persisted++ + } + cp.duplicates += duplicates + cp.mu.Unlock() + w.Header().Set("Content-Type", "application/json") + _ = json.NewEncoder(w).Encode(map[string]int{"persisted": persisted, "duplicates": duplicates}) +} + +func (cp *tokenGatedControlPlane) setValid(token string) { + cp.mu.Lock() + defer cp.mu.Unlock() + cp.valid[token] = true +} + +func (cp *tokenGatedControlPlane) requestTokens() []string { + cp.mu.Lock() + defer cp.mu.Unlock() + return append([]string(nil), cp.tokens...) +} + +func (cp *tokenGatedControlPlane) persistedIDs() map[string]int { + cp.mu.Lock() + defer cp.mu.Unlock() + out := make(map[string]int, len(cp.persisted)) + for id, n := range cp.persisted { + out[id] = n + } + return out +} + +// newHeldTestReporter builds a reporter whose background loop never ticks during +// the test, so every flush below is one the test drives. +func newHeldTestReporter(t *testing.T, endpoint, token string, adjust func(*Config)) (*Reporter, func() int) { + t.Helper() + db := openTestDB(t) + cfg := testConfig(endpoint, "ws-1") + cfg.BatchMaxWait = time.Hour + if adjust != nil { + adjust(&cfg) + } + r, err := New(db, cfg) + if err != nil { + t.Fatalf("new: %v", err) + } + t.Cleanup(r.Shutdown) + r.SetToken(token) + outbox := func() int { + var n int + if err := db.QueryRow("SELECT COUNT(*) FROM message_outbox").Scan(&n); err != nil { + t.Fatalf("count outbox: %v", err) + } + return n + } + return r, outbox +} + +func enqueueAssistant(t *testing.T, r *Reporter, ids ...string) { + t.Helper() + for _, id := range ids { + if err := r.Enqueue(Message{MessageID: id, Role: "assistant", Content: "reply " + id}); err != nil { + t.Fatalf("enqueue %s: %v", id, err) + } + } +} + +func TestCredentialRejection_HoldsRowsAndStopsSending(t *testing.T) { + cp, server := newTokenGatedControlPlane(t) + r, outbox := newHeldTestReporter(t, server.URL, "expired", nil) + + enqueueAssistant(t, r, "m0", "m1") + r.flush() + r.flush() + r.flush() + + if got := cp.requestTokens(); len(got) != 1 { + t.Fatalf("a rejected token must be sent once, then held; requests = %v", got) + } + if got := outbox(); got != 2 { + t.Fatalf("rejected rows must stay queued, outbox = %d", got) + } + // Liveness: the reporter still accepts new messages while held. + enqueueAssistant(t, r, "m2") + if got := outbox(); got != 3 { + t.Fatalf("enqueue while held: outbox = %d, want 3", got) + } +} + +func TestCredentialRejection_ResumesWithReplacementTokenWithoutDuplicates(t *testing.T) { + cp, server := newTokenGatedControlPlane(t, "renewed") + r, outbox := newHeldTestReporter(t, server.URL, "expired", nil) + + enqueueAssistant(t, r, "m0", "m1") + r.flush() + if got := outbox(); got != 2 { + t.Fatalf("outbox after 401 = %d, want 2", got) + } + + r.SetToken("renewed") + r.flush() + + if got := outbox(); got != 0 { + t.Fatalf("outbox after resume = %d, want 0", got) + } + if got := cp.requestTokens(); strings.Join(got, ",") != "expired,renewed" { + t.Fatalf("request tokens = %v", got) + } + persisted := cp.persistedIDs() + if len(persisted) != 2 || persisted["m0"] != 1 || persisted["m1"] != 1 || cp.duplicates != 0 { + t.Fatalf("each message must be persisted exactly once: %v (duplicates %d)", persisted, cp.duplicates) + } +} + +func TestCredentialRejection_SameTokenDoesNotResume(t *testing.T) { + cp, server := newTokenGatedControlPlane(t) + r, outbox := newHeldTestReporter(t, server.URL, "expired", nil) + + enqueueAssistant(t, r, "m0") + r.flush() + r.SetToken("expired") // e.g. getOrCreateReporter re-syncing an unchanged runtime token + r.flush() + + if got := cp.requestTokens(); len(got) != 1 { + t.Fatalf("re-setting the rejected token must not resend; requests = %v", got) + } + if got := outbox(); got != 1 { + t.Fatalf("outbox = %d, want 1", got) + } +} + +func TestCredentialRejection_StaleTokenAfterRotationRetriesImmediately(t *testing.T) { + cp, server := newTokenGatedControlPlane(t, "t2") + r, outbox := newHeldTestReporter(t, server.URL, "t1", nil) + + inFlight := make(chan struct{}) + release := make(chan struct{}) + var once sync.Once + cp.mu.Lock() + cp.onRequest = func(token string) { + if token == "t1" { + once.Do(func() { close(inFlight) }) + <-release + } + } + cp.mu.Unlock() + + enqueueAssistant(t, r, "m0") + done := make(chan struct{}) + go func() { + r.flush() + close(done) + }() + + <-inFlight + r.SetToken("t2") // renewal lands while the t1 request is in flight + close(release) // ...then the control plane answers t1 with 401 + <-done + + if got := cp.requestTokens(); strings.Join(got, ",") != "t1,t2" { + t.Fatalf("a 401 for a replaced token must be resent with the current one; requests = %v", got) + } + if got := outbox(); got != 0 { + t.Fatalf("outbox = %d, want 0", got) + } + r.mu.Lock() + held := r.credentialWait.rejectedToken + r.mu.Unlock() + if held != "" { + t.Fatalf("a stale 401 must not pause delivery on the current token, held on %q", held) + } +} + +func TestCredentialRejection_SurfacesLongPauseOnceWithoutDeleting(t *testing.T) { + cp, server := newTokenGatedControlPlane(t) + var reports []AuthRenewalWaitExceeded + r, outbox := newHeldTestReporter(t, server.URL, "expired", func(cfg *Config) { + cfg.AuthRenewalWait = time.Hour + cfg.OnAuthRenewalWaitExceeded = func(info AuthRenewalWaitExceeded) { + reports = append(reports, info) + } + }) + clock := time.Date(2026, 10, 4, 8, 0, 0, 0, time.UTC) + r.now = func() time.Time { return clock } + + enqueueAssistant(t, r, "m0", "m1") + r.flush() // 401: pause starts + + clock = clock.Add(30 * time.Minute) + r.flush() + if len(reports) != 0 { + t.Fatalf("pause reported before the budget: %+v", reports) + } + + clock = clock.Add(31 * time.Minute) + r.flush() + r.flush() + if len(reports) != 1 { + t.Fatalf("pause must be reported exactly once, got %d", len(reports)) + } + if got := reports[0]; got.WorkspaceID != "ws-1" || got.SessionID != "sess-1" || got.HeldMessages != 2 || got.PausedFor < time.Hour { + t.Fatalf("unexpected report: %+v", got) + } + if got := outbox(); got != 2 { + t.Fatalf("an exhausted wait must keep the queued rows, outbox = %d", got) + } + + // A replacement token still delivers the held rows. + cp.setValid("renewed") + r.SetToken("renewed") + r.flush() + if got := outbox(); got != 0 { + t.Fatalf("outbox after late replacement = %d, want 0", got) + } + + // A later, separate pause is reported on its own. + r.SetToken("rejected-again") + enqueueAssistant(t, r, "m2") + r.flush() + clock = clock.Add(2 * time.Hour) + r.flush() + if len(reports) != 2 { + t.Fatalf("a new pause must be reported again, got %d reports", len(reports)) + } +} + +func TestCredentialRejection_DuringSizeFallbackKeepsRowsAndResendsOnce(t *testing.T) { + cp, server := newTokenGatedControlPlane(t, "valid") + cp.batchLimit = 1 + r, outbox := newHeldTestReporter(t, server.URL, "valid", nil) + + enqueueAssistant(t, r, "m0", "m1") + // Requests: the batch (400, too large) → m0 alone (200) → m1 alone. The token + // stops being accepted just before m1, so m0 is persisted and m1 is not. + sent := 0 + cp.mu.Lock() + cp.onRequest = func(string) { + cp.mu.Lock() + defer cp.mu.Unlock() + sent++ + if sent == 3 { + delete(cp.valid, "valid") + } + } + cp.mu.Unlock() + + r.flush() + if got := outbox(); got != 2 { + t.Fatalf("a 401 during the row-by-row fallback must keep the batch, outbox = %d", got) + } + + cp.setValid("renewed") + r.SetToken("renewed") + r.flush() + + if got := outbox(); got != 0 { + t.Fatalf("outbox after resume = %d, want 0", got) + } + persisted := cp.persistedIDs() + if persisted["m0"] != 1 || persisted["m1"] != 1 || len(persisted) != 2 { + t.Fatalf("each message must be persisted exactly once: %v", persisted) + } + if cp.duplicates != 1 { + t.Fatalf("the resent m0 must be absorbed as a duplicate, duplicates = %d", cp.duplicates) + } +} + +func TestCredentialRejection_HeldOutboxStaysBounded(t *testing.T) { + _, server := newTokenGatedControlPlane(t) + r, outbox := newHeldTestReporter(t, server.URL, "expired", func(cfg *Config) { + cfg.OutboxMaxSize = 2 + }) + + enqueueAssistant(t, r, "m0") + r.flush() + enqueueAssistant(t, r, "m1") + if err := r.Enqueue(Message{MessageID: "m2", Role: "assistant", Content: "over"}); err == nil { + t.Fatal("enqueue beyond the outbox cap must fail explicitly while held") + } + if got := outbox(); got != 2 { + t.Fatalf("outbox = %d, want 2", got) + } +} diff --git a/packages/vm-agent/internal/messagereport/reporter.go b/packages/vm-agent/internal/messagereport/reporter.go index a6c2c27d68..ab9eae3827 100644 --- a/packages/vm-agent/internal/messagereport/reporter.go +++ b/packages/vm-agent/internal/messagereport/reporter.go @@ -3,6 +3,7 @@ package messagereport import ( "context" "database/sql" + "errors" "fmt" "log/slog" "net/http" @@ -48,6 +49,10 @@ type Reporter struct { terminalPersistenceFailure bool terminalPersistenceReason string terminalWakeC chan struct{} + // credentialWait pauses delivery while the control plane rejects the + // current token (see credential.go). Guarded by mu. + credentialWait credentialWait + now func() time.Time // flushMu serializes flush() calls with outbox mutations in SetSessionID. // Lock ordering: flushMu must always be acquired BEFORE mu when both @@ -106,6 +111,9 @@ func New(db *sql.DB, cfg Config) (*Reporter, error) { if cfg.ResponseMaxBytes <= 0 { cfg.ResponseMaxBytes = defaults.ResponseMaxBytes } + if cfg.AuthRenewalWait <= 0 { + cfg.AuthRenewalWait = defaults.AuthRenewalWait + } if err := migrateOutbox(db); err != nil { return nil, fmt.Errorf("messagereport: migrate outbox: %w", err) @@ -120,6 +128,7 @@ func New(db *sql.DB, cfg Config) (*Reporter, error) { workspaceID: cfg.WorkspaceID, sessionID: cfg.SessionID, terminalWakeC: make(chan struct{}, 1), + now: time.Now, stopC: make(chan struct{}), stopCtx: stopCtx, stopCancel: stopCancel, @@ -131,14 +140,22 @@ func New(db *sql.DB, cfg Config) (*Reporter, error) { } // SetToken updates the authorization token used for HTTP POSTs. -// Call this after bootstrap when the callback JWT becomes available. +// Call this after bootstrap when the callback JWT becomes available, and whenever +// the workspace token is renewed or re-delivered. A token different from one the +// control plane rejected resumes delivery of the held messages. func (r *Reporter) SetToken(token string) { if r == nil { return } r.mu.Lock() r.authToken = token + resumed := r.credentialWait.resumeWith(token) + wsID, sessionID := r.workspaceID, r.sessionID r.mu.Unlock() + if resumed { + slog.Info("messagereport: workspace callback token replaced, resuming held messages", + "workspaceId", wsID, "sessionId", sessionID) + } } // SetWorkspaceID updates the workspace ID used in the batch POST URL. @@ -328,6 +345,7 @@ func (r *Reporter) flush() { r.flushMu.Lock() defer r.flushMu.Unlock() + rotatedRetries := 0 for { batch, err := r.readBatch() if err != nil { @@ -339,6 +357,15 @@ func (r *Reporter) flush() { } if err := r.sendBatch(batch); err != nil { + if errors.Is(err, errCredentialRotated) && rotatedRetries < maxRotatedTokenRetriesPerFlush { + // A renewed token replaced the one that got 401: resend now. + rotatedRetries++ + continue + } + if errors.Is(err, errAwaitingCredential) { + // Held until a new token arrives; not a delivery attempt. + return + } // sendBatch handles retry internally; if it returns an error the // batch was NOT sent and remains in the outbox for the next tick. slog.Warn("messagereport: send batch failed", "error", err, "count", len(batch)) diff --git a/packages/vm-agent/internal/messagereport/reporter_test.go b/packages/vm-agent/internal/messagereport/reporter_test.go index f5f8e58855..bdb7884edb 100644 --- a/packages/vm-agent/internal/messagereport/reporter_test.go +++ b/packages/vm-agent/internal/messagereport/reporter_test.go @@ -426,9 +426,10 @@ func TestFlush_PermanentError_Discards(t *testing.T) { } } +// 401 is deliberately absent: a rejected token pauses delivery instead +// (TestCredentialRejection_* in credential_test.go). func TestFlush_TerminalStatusesDisableFutureMessages(t *testing.T) { for _, status := range []int{ - http.StatusUnauthorized, http.StatusForbidden, http.StatusNotFound, http.StatusGone, @@ -994,8 +995,8 @@ func TestFlush_SizeFallbackPermanentErrorDiscardsBatch(t *testing.T) { return } if ids[0] == "m1" { - w.WriteHeader(http.StatusUnauthorized) - _, _ = w.Write([]byte(`{"error":"UNAUTHORIZED"}`)) + w.WriteHeader(http.StatusForbidden) + _, _ = w.Write([]byte(`{"error":"FORBIDDEN"}`)) return } diff --git a/packages/vm-agent/internal/messagereport/sender.go b/packages/vm-agent/internal/messagereport/sender.go index d2530e4cd9..c5c88d0764 100644 --- a/packages/vm-agent/internal/messagereport/sender.go +++ b/packages/vm-agent/internal/messagereport/sender.go @@ -82,6 +82,9 @@ func (r *Reporter) sendBatch(batch []outboxRow) error { // No workspace yet — leave messages in outbox for later. return fmt.Errorf("no workspace ID") } + if r.awaitingCredential(token) { + return errAwaitingCredential + } body, err := buildBatchBody(batch) if err != nil { @@ -137,6 +140,9 @@ func (r *Reporter) handleBatchResponse(batch []outboxRow, url, token, wsID strin if postErr == nil && statusCode >= 200 && statusCode < 300 { return true, nil } + if statusCode == http.StatusUnauthorized { + return true, r.credentialRejected(token, responseBody) + } if statusCode == http.StatusConflict && isSessionMessageLimitError(responseBody) { r.markMessageLimitReached(batch, responseBody) return true, nil @@ -205,9 +211,11 @@ func (r *Reporter) discardRejectedRow(row outboxRow, rejection rowRejectedError) r.deleteBatch([]outboxRow{row}) } +// isTerminalBatchResponse is true when the control plane says the workspace or +// session will never accept these messages. 401 is not terminal: it rejects the +// token, which can be replaced (see credential.go). func isTerminalBatchResponse(statusCode int) bool { - return statusCode == http.StatusUnauthorized || - statusCode == http.StatusForbidden || + return statusCode == http.StatusForbidden || statusCode == http.StatusNotFound || statusCode == http.StatusGone } @@ -332,6 +340,10 @@ func (r *Reporter) sendRowsIndividually(url, token, wsID string, batch []outboxR case sessionMessageLimitError: r.markMessageLimitReached(batch, verdict.responseBody) return nil + case credentialRejectedError: + // Rows already sent are resent with the next token; the control plane + // dedupes them by message id. + return r.credentialRejected(token, verdict.responseBody) case terminalPersistenceError: r.markTerminalPersistenceFailure(batch, verdict.statusCode, verdict.responseBody) return nil @@ -392,6 +404,9 @@ func fallbackCandidateResult(row outboxRow, candidateIndex int, statusCode int, if statusCode == http.StatusConflict && isSessionMessageLimitError(responseBody) { return false, sessionMessageLimitError{responseBody: responseBody} } + if statusCode == http.StatusUnauthorized { + return false, credentialRejectedError{responseBody: responseBody} + } if isTerminalBatchResponse(statusCode) { return false, terminalPersistenceError{statusCode: statusCode, responseBody: responseBody} } diff --git a/packages/vm-agent/internal/server/health.go b/packages/vm-agent/internal/server/health.go index 0a2ea8a1bf..e3a65ce047 100644 --- a/packages/vm-agent/internal/server/health.go +++ b/packages/vm-agent/internal/server/health.go @@ -378,6 +378,8 @@ func (s *Server) sendNodeHeartbeat() { // Heartbeat succeeded — connectivity to the control plane is confirmed. // Resume one durable eviction callback without delaying the heartbeat ticker. go s.retryPendingEvictionCallbacks() + // Renew workspace callback tokens that are past their refresh point. + go s.renewDueWorkspaceCallbackTokensOnce() // Retry any pending workspace-ready callbacks in a background goroutine // so the heartbeat ticker is not blocked by potentially slow HTTP calls. diff --git a/packages/vm-agent/internal/server/server.go b/packages/vm-agent/internal/server/server.go index 22b338831c..c1b6f4711f 100644 --- a/packages/vm-agent/internal/server/server.go +++ b/packages/vm-agent/internal/server/server.go @@ -123,6 +123,7 @@ type Server struct { bootstrapComplete atomic.Bool callbackTokenMu sync.RWMutex callbackToken string + tokenRenewal workspaceTokenRenewal // workspace callback token renewal (workspace_callback_token_renewal.go) callbacksTerminal atomic.Bool httpClient *http.Client // shared HTTP client with timeout for control-plane callbacks done chan struct{} @@ -1292,6 +1293,7 @@ func (s *Server) getOrCreateReporter(workspaceID, projectID, chatSessionID strin // Slow path: create reporter outside the lock (disk I/O). cfg := messagereport.LoadConfigFromEnv() + cfg.OnAuthRenewalWaitExceeded = s.reportMessagePersistencePaused cfg.ProjectID = projectID cfg.SessionID = chatSessionID cfg.WorkspaceID = workspaceID diff --git a/packages/vm-agent/internal/server/workspace_callback_token_renewal.go b/packages/vm-agent/internal/server/workspace_callback_token_renewal.go new file mode 100644 index 0000000000..8b3c2c6b3f --- /dev/null +++ b/packages/vm-agent/internal/server/workspace_callback_token_renewal.go @@ -0,0 +1,407 @@ +package server + +// Workspace callback token renewal and propagation. +// +// The control plane mints each workspace a workspace-scoped callback token +// (CALLBACK_TOKEN_EXPIRY_MS, default 24h) when it creates or restores the +// workspace. Every workspace callback authenticates with it: messages, snapshot +// prepare/progress/complete/failure, git-token, runtime-assets, task status, +// ACP activity/usage/interactions and more. A workspace awake longer than the +// lifetime used to fail all of them with 401. +// +// Two paths keep runtime.CallbackToken fresh, and every change is persisted and +// then propagated to the consumers that copied it (message reporter, SessionHosts): +// +// 1. Renewal. After each successful node heartbeat, every token past +// WorkspaceCallbackTokenRefreshRatio of its lifetime is renewed through +// POST /api/workspaces/:id/callback-token/renew, presenting the current +// workspace token AND this node's token. The control plane renews only for +// the node it binds the workspace to, and never renews an expired token. +// 2. Control-plane delivery. Requests the control plane sends over the +// node-management channel (create, restore, and VM hibernate) carry a fresh +// token; upsertWorkspaceRuntime adopts it (adoptWorkspaceCallbackTokenLocked). + +import ( + "bytes" + "context" + "encoding/json" + "errors" + "io" + "log/slog" + "net/http" + neturl "net/url" + "sort" + "strings" + "sync" + "time" + + "github.com/golang-jwt/jwt/v5" + + "github.com/workspace/vm-agent/internal/messagereport" +) + +// renewalResponseMaxBytes bounds how much of a renewal response is read. +const renewalResponseMaxBytes = 16 * 1024 + +var errMessagePersistencePaused = errors.New( + "chat message persistence paused: the control plane rejected the workspace callback token and no replacement arrived") + +// Node-credential error codes from the control plane. The node token refreshes +// through the heartbeat, so these are retried; they say nothing about the +// workspace token (apps/api/src/services/workspace-callback-token-renewal.ts). +const ( + renewalNodeCallbackUnauthorized = "NODE_CALLBACK_UNAUTHORIZED" + renewalNodeCallbackForbidden = "NODE_CALLBACK_FORBIDDEN" +) + +// workspaceTokenRenewal is the renewal bookkeeping on Server. +type workspaceTokenRenewal struct { + running sync.Mutex // one renewal pass at a time (TryLock from the heartbeat) + mu sync.Mutex + state map[string]workspaceTokenRenewalState // keyed by workspace ID + now func() time.Time // test clock; nil means time.Now +} + +// workspaceTokenRenewalState is keyed to the token it describes: when the +// workspace's token changes (renewal or control-plane delivery) the state no +// longer applies, so a new token is never blocked by an old token's latch or +// backoff. +type workspaceTokenRenewalState struct { + token string + rejected bool // the control plane refused this token; never present it again + nextAttempt time.Time // earliest next attempt after a transient failure + failures int // consecutive transient failures, for backoff +} + +type workspaceTokenRenewalCandidate struct { + workspaceID string + token string +} + +type workspaceTokenRenewalOutcome int + +const ( + renewalSucceeded workspaceTokenRenewalOutcome = iota + renewalNotDue + renewalRejected + renewalTransientFailure +) + +func (s *Server) workspaceTokenRenewalNow() time.Time { + if now := s.tokenRenewal.now; now != nil { + return now() + } + return time.Now() +} + +// renewDueWorkspaceCallbackTokensOnce runs one renewal pass unless one is +// already running. Called in its own goroutine after a successful heartbeat. +func (s *Server) renewDueWorkspaceCallbackTokensOnce() { + if !s.tokenRenewal.running.TryLock() { + return + } + defer s.tokenRenewal.running.Unlock() + s.renewDueWorkspaceCallbackTokens() +} + +func (s *Server) renewDueWorkspaceCallbackTokens() { + if s.controlPlaneCallbacksStopped() || s.config == nil || + s.config.ControlPlaneURL == "" || s.config.NodeID == "" { + return + } + nodeToken := strings.TrimSpace(s.getCallbackToken()) + if nodeToken == "" { + return + } + for _, candidate := range s.dueWorkspaceCallbackTokenRenewals(s.workspaceTokenRenewalNow()) { + s.renewWorkspaceCallbackToken(candidate, nodeToken) + } +} + +// dueWorkspaceCallbackTokenRenewals lists workspaces whose current token is due: +// past the refresh ratio, not refused by the control plane, and not backing off. +// An expired token is still offered once; the control plane, not this clock, +// decides it is expired, so agent clock skew cannot strand a valid token. +func (s *Server) dueWorkspaceCallbackTokenRenewals(now time.Time) []workspaceTokenRenewalCandidate { + s.workspaceMu.RLock() + tokens := make(map[string]string, len(s.workspaces)) + for id, runtime := range s.workspaces { + if token := strings.TrimSpace(runtime.CallbackToken); token != "" { + tokens[id] = token + } + } + s.workspaceMu.RUnlock() + + ratio := s.config.WorkspaceCallbackTokenRefreshRatio + s.tokenRenewal.mu.Lock() + defer s.tokenRenewal.mu.Unlock() + if s.tokenRenewal.state == nil { + s.tokenRenewal.state = make(map[string]workspaceTokenRenewalState) + } + for id, state := range s.tokenRenewal.state { + if tokens[id] != state.token { + delete(s.tokenRenewal.state, id) // workspace gone or token replaced + } + } + due := make([]workspaceTokenRenewalCandidate, 0, len(tokens)) + for id, token := range tokens { + state := s.tokenRenewal.state[id] + if state.rejected || now.Before(state.nextAttempt) { + continue + } + if callbackTokenRenewalDue(token, now, ratio) { + due = append(due, workspaceTokenRenewalCandidate{workspaceID: id, token: token}) + } + } + sort.Slice(due, func(i, j int) bool { return due[i].workspaceID < due[j].workspaceID }) + return due +} + +// callbackTokenRenewalDue reports whether now is at least ratio of the way +// through the token's iat→exp lifetime. A token whose claims cannot be read is +// offered once so the control plane can classify it. +func callbackTokenRenewalDue(token string, now time.Time, ratio float64) bool { + issuedAt, expiresAt, ok := callbackTokenLifetime(token) + if !ok { + return true + } + lifetime := expiresAt.Sub(issuedAt) + return !now.Before(issuedAt.Add(time.Duration(float64(lifetime) * ratio))) +} + +// callbackTokenLifetime reads iat/exp WITHOUT verifying the signature. It only +// schedules renewal; the control plane verifies every token it receives. +func callbackTokenLifetime(token string) (issuedAt, expiresAt time.Time, ok bool) { + claims := jwt.RegisteredClaims{} + if _, _, err := jwt.NewParser().ParseUnverified(token, &claims); err != nil { + return time.Time{}, time.Time{}, false + } + if claims.IssuedAt == nil || claims.ExpiresAt == nil || !claims.ExpiresAt.After(claims.IssuedAt.Time) { + return time.Time{}, time.Time{}, false + } + return claims.IssuedAt.Time, claims.ExpiresAt.Time, true +} + +func (s *Server) renewWorkspaceCallbackToken(candidate workspaceTokenRenewalCandidate, nodeToken string) { + outcome, renewed, detail := s.requestWorkspaceCallbackTokenRenewal(candidate, nodeToken) + now := s.workspaceTokenRenewalNow() + switch outcome { + case renewalSucceeded: + if s.replaceRenewedWorkspaceCallbackToken(candidate.workspaceID, candidate.token, renewed) { + slog.Info("Workspace callback token renewed", "workspace", candidate.workspaceID) + } else { + slog.Info("Discarded renewed workspace callback token; the workspace token changed during renewal", + "workspace", candidate.workspaceID) + } + s.setWorkspaceTokenRenewalState(candidate.workspaceID, workspaceTokenRenewalState{}) + case renewalNotDue: + s.setWorkspaceTokenRenewalState(candidate.workspaceID, workspaceTokenRenewalState{ + token: candidate.token, + nextAttempt: now.Add(s.config.WorkspaceCallbackTokenRenewalRetryMax), + }) + case renewalRejected: + slog.Warn("Control plane refused workspace callback token renewal; waiting for a new token", + "workspace", candidate.workspaceID, "detail", detail) + s.setWorkspaceTokenRenewalState(candidate.workspaceID, workspaceTokenRenewalState{ + token: candidate.token, + rejected: true, + }) + default: + failures := s.workspaceTokenRenewalFailures(candidate.workspaceID, candidate.token) + 1 + delay := workspaceTokenRenewalBackoff(failures, + s.config.WorkspaceCallbackTokenRenewalRetryInitial, s.config.WorkspaceCallbackTokenRenewalRetryMax) + slog.Warn("Workspace callback token renewal failed; will retry", + "workspace", candidate.workspaceID, "detail", detail, "retryIn", delay.String()) + s.setWorkspaceTokenRenewalState(candidate.workspaceID, workspaceTokenRenewalState{ + token: candidate.token, + failures: failures, + nextAttempt: now.Add(delay), + }) + } +} + +func workspaceTokenRenewalBackoff(failures int, initial, max time.Duration) time.Duration { + delay := initial + for i := 1; i < failures && delay < max; i++ { + delay *= 2 + } + if delay > max { + return max + } + return delay +} + +func (s *Server) workspaceTokenRenewalFailures(workspaceID, token string) int { + s.tokenRenewal.mu.Lock() + defer s.tokenRenewal.mu.Unlock() + if state, ok := s.tokenRenewal.state[workspaceID]; ok && state.token == token { + return state.failures + } + return 0 +} + +func (s *Server) setWorkspaceTokenRenewalState(workspaceID string, state workspaceTokenRenewalState) { + s.tokenRenewal.mu.Lock() + defer s.tokenRenewal.mu.Unlock() + if s.tokenRenewal.state == nil { + s.tokenRenewal.state = make(map[string]workspaceTokenRenewalState) + } + if state.token == "" { + delete(s.tokenRenewal.state, workspaceID) + return + } + s.tokenRenewal.state[workspaceID] = state +} + +// requestWorkspaceCallbackTokenRenewal performs one renewal request. detail is a +// status/code summary for logs and never contains a token. +func (s *Server) requestWorkspaceCallbackTokenRenewal( + candidate workspaceTokenRenewalCandidate, + nodeToken string, +) (outcome workspaceTokenRenewalOutcome, renewed string, detail string) { + endpoint := strings.TrimRight(s.config.ControlPlaneURL, "/") + + "/api/workspaces/" + neturl.PathEscape(candidate.workspaceID) + "/callback-token/renew" + body, err := json.Marshal(map[string]string{"nodeId": s.config.NodeID, "nodeToken": nodeToken}) + if err != nil { + return renewalTransientFailure, "", "marshal request: " + err.Error() + } + timeout := s.config.WorkspaceCallbackTokenRenewalTimeout + ctx, cancel := context.WithTimeout(context.Background(), timeout) + defer cancel() + req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, bytes.NewReader(body)) + if err != nil { + return renewalTransientFailure, "", "build request: " + err.Error() + } + req.Header.Set("Authorization", "Bearer "+candidate.token) + req.Header.Set("Content-Type", "application/json") + + resp, err := s.controlPlaneHTTPClient(timeout).Do(req) + if err != nil { + return renewalTransientFailure, "", "request failed: " + err.Error() + } + defer resp.Body.Close() + payload, _ := io.ReadAll(io.LimitReader(resp.Body, renewalResponseMaxBytes)) + + var parsed struct { + Renewed bool `json:"renewed"` + Token string `json:"token"` + Error string `json:"error"` + } + parseErr := json.Unmarshal(payload, &parsed) + detail = "status " + resp.Status + if parsed.Error != "" { + detail += " " + parsed.Error + } + return classifyWorkspaceTokenRenewal(resp.StatusCode, parsed.Renewed, parsed.Token, parsed.Error, parseErr), strings.TrimSpace(parsed.Token), detail +} + +func classifyWorkspaceTokenRenewal(status int, renewed bool, token, errorCode string, parseErr error) workspaceTokenRenewalOutcome { + switch { + case status == http.StatusOK && parseErr == nil && renewed && strings.TrimSpace(token) != "": + return renewalSucceeded + case status == http.StatusOK && parseErr == nil && !renewed: + return renewalNotDue + case status == http.StatusOK: + return renewalTransientFailure // malformed success body + case errorCode == renewalNodeCallbackUnauthorized || errorCode == renewalNodeCallbackForbidden: + return renewalTransientFailure // the node token, not the workspace token, was refused + case status == http.StatusTooManyRequests || status >= 500: + return renewalTransientFailure + case status >= 400: + // 400/401/403/404/410: the workspace token was refused, the workspace is + // gone or no longer bound to this node, or the request can never succeed. + // Stop presenting this token (rule 54.13); a new token resets the latch. + return renewalRejected + default: + return renewalTransientFailure + } +} + +// replaceRenewedWorkspaceCallbackToken installs a renewed token only if the +// runtime still holds the token that was renewed. A token delivered by the +// control plane while the renewal was in flight is at least as fresh and wins. +// The token is persisted before any consumer sees it, so a restart can never +// resume with an older token than a consumer has already used. +func (s *Server) replaceRenewedWorkspaceCallbackToken(workspaceID, renewedFrom, renewed string) bool { + renewed = strings.TrimSpace(renewed) + s.workspaceMu.Lock() + runtime, ok := s.workspaces[workspaceID] + if !ok || renewed == "" || strings.TrimSpace(runtime.CallbackToken) != renewedFrom { + s.workspaceMu.Unlock() + return false + } + runtime.CallbackToken = renewed + runtime.UpdatedAt = nowUTC() + if runtime.Repository != "" && !runtime.MetadataUnavailable { + s.persistWorkspaceMetadata(runtime) + } + s.workspaceMu.Unlock() + s.propagateWorkspaceCallbackToken(workspaceID, renewed) + return true +} + +// adoptWorkspaceCallbackTokenLocked installs a token delivered for an existing +// workspace (create, restore, hibernate) unless it would replace the current +// token with one that expires earlier, so a reordered or late delivery never +// rolls a workspace back to an older credential. Caller holds workspaceMu and +// must persist and then propagate when it returns true. +func adoptWorkspaceCallbackTokenLocked(runtime *WorkspaceRuntime, delivered string) bool { + delivered = strings.TrimSpace(delivered) + current := strings.TrimSpace(runtime.CallbackToken) + if delivered == "" || delivered == current { + return false + } + if current != "" { + _, deliveredExpiry, deliveredOK := callbackTokenLifetime(delivered) + _, currentExpiry, currentOK := callbackTokenLifetime(current) + if deliveredOK && currentOK && deliveredExpiry.Before(currentExpiry) { + slog.Info("Ignoring delivered workspace callback token that expires before the current one", + "workspace", runtime.ID) + return false + } + } + runtime.CallbackToken = delivered + return true +} + +// propagateWorkspaceCallbackToken hands a changed workspace token to every +// consumer that copied the previous one. The message reporter resumes held +// messages (messagereport/credential.go); SessionHosts use it for every later +// control-plane call. Hosts are found under sessionHostMu, which host creation +// also holds while it reads the token, so a host is either created with the new +// token or updated here. Callers must not hold workspaceMu, messageReportersMu +// or sessionHostMu. +func (s *Server) propagateWorkspaceCallbackToken(workspaceID, token string) { + if strings.TrimSpace(token) == "" { + return + } + s.messageReportersMu.RLock() + reporter := s.messageReporters[workspaceID] + s.messageReportersMu.RUnlock() + reporter.SetToken(token) + + prefix := workspaceID + ":" + s.sessionHostMu.Lock() + for key, host := range s.sessionHosts { + if strings.HasPrefix(key, prefix) && host != nil { + host.SetCallbackToken(token) + } + } + s.sessionHostMu.Unlock() +} + +// reportMessagePersistencePaused surfaces a message reporter that has held chat +// messages for longer than MSG_AUTH_RENEWAL_WAIT because the control plane keeps +// rejecting the workspace token. The error reporter authenticates with the node +// token, so the report gets through while the workspace token is unusable. +func (s *Server) reportMessagePersistencePaused(info messagereport.AuthRenewalWaitExceeded) { + if s == nil || s.errorReporter == nil { + return + } + s.errorReporter.ReportError(errMessagePersistencePaused, "messagereport.credential_wait", info.WorkspaceID, + map[string]interface{}{ + "sessionId": info.SessionID, + "heldMessages": info.HeldMessages, + "pausedForSeconds": int64(info.PausedFor / time.Second), + }) +} diff --git a/packages/vm-agent/internal/server/workspace_routing.go b/packages/vm-agent/internal/server/workspace_routing.go index 1b1853a5f8..147b2e5776 100644 --- a/packages/vm-agent/internal/server/workspace_routing.go +++ b/packages/vm-agent/internal/server/workspace_routing.go @@ -209,8 +209,10 @@ func (s *Server) upsertWorkspaceRuntime(workspaceID, repository, branch, status, } s.workspaceMu.Lock() var resourceHistorySnapshot *WorkspaceRuntime + var adoptedCallbackToken string // persisted under the lock, published after it defer func() { s.workspaceMu.Unlock() + s.propagateWorkspaceCallbackToken(workspaceID, adoptedCallbackToken) if resourceHistorySnapshot != nil { s.ensureResourceHistoryForRuntime(resourceHistorySnapshot) } @@ -240,8 +242,8 @@ func (s *Server) upsertWorkspaceRuntime(workspaceID, repository, branch, status, if status != "" && runtime.Status != "evicted" && !runtime.ProvisioningActive && !runtime.MetadataUnavailable { runtime.Status = status } - if callbackToken != "" { - runtime.CallbackToken = strings.TrimSpace(callbackToken) + if adoptWorkspaceCallbackTokenLocked(runtime, callbackToken) { + adoptedCallbackToken, metadataChanged = runtime.CallbackToken, true } if runtime.WorkspaceDir == "" { runtime.WorkspaceDir = s.workspaceDirForRepo(workspaceID, runtime.Repository) @@ -423,6 +425,7 @@ func (s *Server) upsertWorkspaceRuntime(workspaceID, repository, branch, status, PTY: manager, } s.workspaces[workspaceID] = runtime + adoptedCallbackToken = runtime.CallbackToken runtimeCopy := *runtime resourceHistorySnapshot = &runtimeCopy From d3f976a8c044f46dce35b3d836a7a976281c33c5 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 09:02:47 +0000 Subject: [PATCH 05/19] test(vm-agent): cover workspace token renewal, delivery and held messages - Renewal crosses the refresh point and the 24h expiry on an injected clock, latches refusals per token, backs off transient and node-credential failures, and loses the compare-and-swap to a concurrent delivery. - Deliveries never roll a workspace back to an earlier-expiring token, and every change reaches the parked message reporter and each SessionHost of that workspace only. A renewed token survives an agent restart. - The hibernate handler uses a delivered token for prepare/progress/complete; the same test passes against the pinned pre-fix agent source (7a9782c90), with a no-token control that reproduces the production 401. - Reporter: held rows survive 401, a rotation mid-request resends at once, a long pause is surfaced once, resent rows are absorbed as duplicates. Co-Authored-By: Claude Opus 5.5 --- .../acp/session_host_callback_token.go | 7 + .../acp/session_host_callback_token_test.go | 103 ++++ .../config/callback_token_renewal_test.go | 77 +++ .../internal/messagereport/credential_test.go | 4 + ...workspace_callback_token_hibernate_test.go | 197 +++++++ .../workspace_callback_token_renewal_test.go | 533 ++++++++++++++++++ 6 files changed, 921 insertions(+) create mode 100644 packages/vm-agent/internal/acp/session_host_callback_token_test.go create mode 100644 packages/vm-agent/internal/config/callback_token_renewal_test.go create mode 100644 packages/vm-agent/internal/server/workspace_callback_token_hibernate_test.go create mode 100644 packages/vm-agent/internal/server/workspace_callback_token_renewal_test.go diff --git a/packages/vm-agent/internal/acp/session_host_callback_token.go b/packages/vm-agent/internal/acp/session_host_callback_token.go index 817c6ef506..f5593d1dac 100644 --- a/packages/vm-agent/internal/acp/session_host_callback_token.go +++ b/packages/vm-agent/internal/acp/session_host_callback_token.go @@ -26,3 +26,10 @@ func (h *SessionHost) SetCallbackToken(token string) { h.renewedCallbackToken.Store(token) } } + +// UsesCallbackToken reports whether this host's control-plane calls currently +// authenticate with token, without exposing the token itself. Used by the +// server package to verify that renewals reach every live host. +func (h *SessionHost) UsesCallbackToken(token string) bool { + return token != "" && h.callbackToken() == token +} diff --git a/packages/vm-agent/internal/acp/session_host_callback_token_test.go b/packages/vm-agent/internal/acp/session_host_callback_token_test.go new file mode 100644 index 0000000000..6853e09548 --- /dev/null +++ b/packages/vm-agent/internal/acp/session_host_callback_token_test.go @@ -0,0 +1,103 @@ +package acp + +import ( + "net/http" + "net/http/httptest" + "strings" + "sync" + "testing" + "time" +) + +// activityAuthorizations records the bearer token of every activity report. +func activityAuthorizations(t *testing.T) (*httptest.Server, func() []string) { + t.Helper() + var mu sync.Mutex + var tokens []string + server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + if strings.HasSuffix(r.URL.Path, "/activity") { + mu.Lock() + tokens = append(tokens, strings.TrimPrefix(r.Header.Get("Authorization"), "Bearer ")) + mu.Unlock() + } + w.WriteHeader(http.StatusNoContent) + })) + t.Cleanup(server.Close) + return server, func() []string { + mu.Lock() + defer mu.Unlock() + return append([]string(nil), tokens...) + } +} + +func TestSessionHostControlPlaneCallsUseRenewedCallbackToken(t *testing.T) { + server, seen := activityAuthorizations(t) + host := NewSessionHost(SessionHostConfig{GatewayConfig: GatewayConfig{ + ProjectID: "project", NodeID: "node", SessionID: "session", + ControlPlaneURL: server.URL, CallbackToken: "created-with", HTTPClient: server.Client(), + }}) + defer host.Stop() + + host.reportActivity("idle") + waitFor(t, time.Second, func() bool { return len(seen()) == 1 }) + + host.SetCallbackToken("renewed") + host.reportActivity("idle") + waitFor(t, time.Second, func() bool { return len(seen()) == 2 }) + + if got := strings.Join(seen(), ","); got != "created-with,renewed" { + t.Fatalf("activity reports authenticated with %q, want created-with then renewed", got) + } + if !host.UsesCallbackToken("renewed") || host.UsesCallbackToken("created-with") { + t.Fatal("UsesCallbackToken disagrees with the token the host sends") + } +} + +func TestSessionHostIgnoresEmptyCallbackToken(t *testing.T) { + host := NewSessionHost(SessionHostConfig{GatewayConfig: GatewayConfig{CallbackToken: "created-with"}}) + defer host.Stop() + + host.SetCallbackToken(" ") + if !host.UsesCallbackToken("created-with") { + t.Fatal("an empty delivery must not clear the host's callback token") + } + if host.UsesCallbackToken("") { + t.Fatal("UsesCallbackToken must never match an empty token") + } +} + +// Renewals arrive on the heartbeat goroutine while control-plane calls read the +// token on others, including the ACP notification goroutine (rule 46). Run under +// -race: the accessor must be safe without h.mu. +func TestSessionHostCallbackTokenIsRaceFreeAcrossGoroutines(t *testing.T) { + host := NewSessionHost(SessionHostConfig{GatewayConfig: GatewayConfig{CallbackToken: "t0"}}) + defer host.Stop() + + var wg sync.WaitGroup + stop := make(chan struct{}) + reads := 0 + wg.Add(1) + go func() { + defer wg.Done() + for { + select { + case <-stop: + return + default: + if token := host.callbackToken(); token == "" { + t.Error("reader observed an empty callback token") + return + } + reads++ + } + } + }() + for i := 0; i < 2000; i++ { + host.SetCallbackToken("t" + strings.Repeat("x", i%7+1)) + } + close(stop) + wg.Wait() + if reads == 0 { + t.Fatal("reader goroutine never ran; the race check proved nothing") + } +} diff --git a/packages/vm-agent/internal/config/callback_token_renewal_test.go b/packages/vm-agent/internal/config/callback_token_renewal_test.go new file mode 100644 index 0000000000..147c843ca3 --- /dev/null +++ b/packages/vm-agent/internal/config/callback_token_renewal_test.go @@ -0,0 +1,77 @@ +package config + +import ( + "testing" + "time" +) + +func loadRenewalConfig(t *testing.T, env map[string]string) *Config { + t.Helper() + t.Setenv("CONTROL_PLANE_URL", "https://api.example.com") + t.Setenv("WORKSPACE_ID", "ws-123") + for _, key := range []string{ + "WORKSPACE_CALLBACK_TOKEN_REFRESH_RATIO", + "WORKSPACE_CALLBACK_TOKEN_RENEWAL_TIMEOUT", + "WORKSPACE_CALLBACK_TOKEN_RENEWAL_RETRY_INITIAL", + "WORKSPACE_CALLBACK_TOKEN_RENEWAL_RETRY_MAX", + } { + t.Setenv(key, env[key]) + } + cfg, err := Load() + if err != nil { + t.Fatalf("Load returned error: %v", err) + } + return cfg +} + +func TestWorkspaceCallbackTokenRenewalDefaults(t *testing.T) { + cfg := loadRenewalConfig(t, nil) + + if cfg.WorkspaceCallbackTokenRefreshRatio != DefaultWorkspaceCallbackTokenRefreshRatio || + cfg.WorkspaceCallbackTokenRenewalTimeout != DefaultWorkspaceCallbackTokenRenewalTimeout || + cfg.WorkspaceCallbackTokenRenewalRetryInitial != DefaultWorkspaceCallbackTokenRenewalRetryInitial || + cfg.WorkspaceCallbackTokenRenewalRetryMax != DefaultWorkspaceCallbackTokenRenewalRetryMax { + t.Fatalf("unexpected defaults: ratio=%v timeout=%v initial=%v max=%v", + cfg.WorkspaceCallbackTokenRefreshRatio, cfg.WorkspaceCallbackTokenRenewalTimeout, + cfg.WorkspaceCallbackTokenRenewalRetryInitial, cfg.WorkspaceCallbackTokenRenewalRetryMax) + } +} + +func TestWorkspaceCallbackTokenRefreshRatioIsClamped(t *testing.T) { + for _, tc := range []struct { + value string + want float64 + }{ + {"0.7", 0.7}, + {"0.95", MaxWorkspaceCallbackTokenRefreshRatio}, + {"0.01", MinWorkspaceCallbackTokenRefreshRatio}, + {"0", DefaultWorkspaceCallbackTokenRefreshRatio}, + {"-1", DefaultWorkspaceCallbackTokenRefreshRatio}, + {"not-a-number", DefaultWorkspaceCallbackTokenRefreshRatio}, + {"NaN", DefaultWorkspaceCallbackTokenRefreshRatio}, + } { + t.Run(tc.value, func(t *testing.T) { + cfg := loadRenewalConfig(t, map[string]string{"WORKSPACE_CALLBACK_TOKEN_REFRESH_RATIO": tc.value}) + if cfg.WorkspaceCallbackTokenRefreshRatio != tc.want { + t.Fatalf("ratio %q loaded as %v, want %v", tc.value, cfg.WorkspaceCallbackTokenRefreshRatio, tc.want) + } + }) + } +} + +func TestWorkspaceCallbackTokenRenewalDurationsRejectNonPositiveValues(t *testing.T) { + cfg := loadRenewalConfig(t, map[string]string{ + "WORKSPACE_CALLBACK_TOKEN_RENEWAL_TIMEOUT": "0s", + "WORKSPACE_CALLBACK_TOKEN_RENEWAL_RETRY_INITIAL": "-5m", + "WORKSPACE_CALLBACK_TOKEN_RENEWAL_RETRY_MAX": "2h", + }) + if cfg.WorkspaceCallbackTokenRenewalTimeout != DefaultWorkspaceCallbackTokenRenewalTimeout { + t.Fatalf("timeout = %v, want default", cfg.WorkspaceCallbackTokenRenewalTimeout) + } + if cfg.WorkspaceCallbackTokenRenewalRetryInitial != DefaultWorkspaceCallbackTokenRenewalRetryInitial { + t.Fatalf("retry initial = %v, want default", cfg.WorkspaceCallbackTokenRenewalRetryInitial) + } + if cfg.WorkspaceCallbackTokenRenewalRetryMax != 2*time.Hour { + t.Fatalf("retry max = %v, want override", cfg.WorkspaceCallbackTokenRenewalRetryMax) + } +} diff --git a/packages/vm-agent/internal/messagereport/credential_test.go b/packages/vm-agent/internal/messagereport/credential_test.go index 5162642570..18bbffc91e 100644 --- a/packages/vm-agent/internal/messagereport/credential_test.go +++ b/packages/vm-agent/internal/messagereport/credential_test.go @@ -315,6 +315,10 @@ func TestCredentialRejection_DuringSizeFallbackKeepsRowsAndResendsOnce(t *testin if got := outbox(); got != 2 { t.Fatalf("a 401 during the row-by-row fallback must keep the batch, outbox = %d", got) } + r.flush() + if got := len(cp.requestTokens()); got != 3 { + t.Fatalf("a 401 during the fallback must pause delivery like a batch 401; requests = %d, want 3", got) + } cp.setValid("renewed") r.SetToken("renewed") diff --git a/packages/vm-agent/internal/server/workspace_callback_token_hibernate_test.go b/packages/vm-agent/internal/server/workspace_callback_token_hibernate_test.go new file mode 100644 index 0000000000..d805b0f2d8 --- /dev/null +++ b/packages/vm-agent/internal/server/workspace_callback_token_hibernate_test.go @@ -0,0 +1,197 @@ +package server + +// These tests drive the hibernate handler exactly as the control plane does after +// apps/api/src/services/node-agent-session-snapshots.ts started delivering a fresh +// workspace token with every VM hibernate request. They only use the handler, the +// node-management JWT helpers and a stub control plane, so the same file also runs +// against older VM-agent sources (the agents still on running nodes) unchanged. + +import ( + "bytes" + "net/http" + "net/http/httptest" + "os" + "os/exec" + "path/filepath" + "strings" + "sync" + "testing" + "time" + + "github.com/workspace/vm-agent/internal/config" +) + +const ( + hibernateExpiredToken = "workspace-token-issued-at-create-and-now-expired" + hibernateFreshToken = "workspace-token-delivered-with-hibernate" +) + +// snapshotCallbackRecorder is a control plane that accepts only the fresh token +// and records which token every snapshot callback presented. +type snapshotCallbackRecorder struct { + mu sync.Mutex + tokens map[string][]string // callback suffix → bearer tokens + done chan struct{} + once sync.Once +} + +func (rec *snapshotCallbackRecorder) serve(w http.ResponseWriter, r *http.Request) { + token := strings.TrimPrefix(r.Header.Get("Authorization"), "Bearer ") + suffix := r.URL.Path + if i := strings.Index(suffix, "/session-snapshot/"); i >= 0 { + suffix = suffix[i:] + } + rec.mu.Lock() + rec.tokens[suffix] = append(rec.tokens[suffix], token) + rec.mu.Unlock() + finish := func() { rec.once.Do(func() { close(rec.done) }) } + + if token != hibernateFreshToken { + // A refused prepare ends the capture before a generation exists, so no + // failure callback follows; a refused failure report ends it too. + if strings.HasSuffix(suffix, "/prepare") || strings.HasSuffix(suffix, "/failure") { + finish() + } + w.Header().Set("Content-Type", "application/json") + w.WriteHeader(http.StatusUnauthorized) + _, _ = w.Write([]byte(`{"error":"UNAUTHORIZED","message":"Invalid or expired callback token"}`)) + return + } + switch { + case strings.HasSuffix(suffix, "/session-snapshot/prepare"): + w.Header().Set("Content-Type", "application/json") + _, _ = w.Write([]byte(`{"generation":"01GENERATION","config":{"totalBudgetBytes":67108864,"entryThresholdBytes":1048576,"transferIdleTimeoutMs":30000,"jsonBodyMaxBytes":262144},"upload":{"home":"/upload/home","wip":"/upload/wip"}}`)) + case strings.HasSuffix(suffix, "/session-snapshot/complete"): + w.Header().Set("Content-Type", "application/json") + _, _ = w.Write([]byte(`{"status":"available"}`)) + finish() + case strings.HasSuffix(suffix, "/session-snapshot/failure"): + finish() + w.WriteHeader(http.StatusNoContent) + default: + w.WriteHeader(http.StatusNoContent) + } +} + +func (rec *snapshotCallbackRecorder) seen() map[string][]string { + rec.mu.Lock() + defer rec.mu.Unlock() + out := make(map[string][]string, len(rec.tokens)) + for suffix, tokens := range rec.tokens { + out[suffix] = append([]string(nil), tokens...) + } + return out +} + +func hibernateRunGit(t *testing.T, dir string, args ...string) { + t.Helper() + cmd := exec.Command("git", args...) + cmd.Dir = dir + cmd.Env = append(os.Environ(), "GIT_AUTHOR_NAME=t", "GIT_AUTHOR_EMAIL=t@example.com", + "GIT_COMMITTER_NAME=t", "GIT_COMMITTER_EMAIL=t@example.com") + if out, err := cmd.CombinedOutput(); err != nil { + t.Fatalf("git %v: %v\n%s", args, err, out) + } +} + +// hibernateTestRepo is a clone of a bare remote with one pushed commit and +// uncommitted local work, the smallest workspace a capture can snapshot. +func hibernateTestRepo(t *testing.T) string { + t.Helper() + root := t.TempDir() + remote := filepath.Join(root, "remote.git") + hibernateRunGit(t, root, "init", "--bare", "--initial-branch=main", remote) + repo := filepath.Join(root, "repo") + hibernateRunGit(t, root, "clone", remote, repo) + if err := os.WriteFile(filepath.Join(repo, "README.md"), []byte("hello\n"), 0o644); err != nil { + t.Fatal(err) + } + hibernateRunGit(t, repo, "add", "README.md") + hibernateRunGit(t, repo, "commit", "-m", "init") + hibernateRunGit(t, repo, "push", "origin", "HEAD:main") + if err := os.WriteFile(filepath.Join(repo, "README.md"), []byte("hello\nlocal work\n"), 0o644); err != nil { + t.Fatal(err) + } + return repo +} + +// hibernateWithBody calls the real hibernate handler the way the control plane +// does (node-management JWT, background capture) and waits for the capture to +// report completion or failure. +func hibernateWithBody(t *testing.T, body string) map[string][]string { + t.Helper() + home := t.TempDir() + if err := os.WriteFile(filepath.Join(home, ".gitconfig"), []byte("[user]\n\tname = t\n"), 0o644); err != nil { + t.Fatal(err) + } + t.Setenv("HOME", home) + + rec := &snapshotCallbackRecorder{tokens: map[string][]string{}, done: make(chan struct{})} + controlPlane := httptest.NewServer(http.HandlerFunc(rec.serve)) + defer controlPlane.Close() + + validator, key := newWorkspaceCreateJWTValidator(t, "node-1") + s := &Server{ + config: &config.Config{ + Role: config.RoleStandalone, + NodeID: "node-1", + ControlPlaneURL: controlPlane.URL, + SessionSnapshotOperationTimeout: time.Minute, + SessionSnapshotProgressReportTimeout: 5 * time.Second, + }, + jwtValidator: validator, + workspaces: map[string]*WorkspaceRuntime{ + "ws-1": {ID: "ws-1", CallbackToken: hibernateExpiredToken, WorkspaceDir: hibernateTestRepo(t), Status: "running"}, + }, + } + + req := httptest.NewRequest(http.MethodPost, "/workspaces/ws-1/agent-sessions/agent-1/hibernate", bytes.NewBufferString(body)) + req.SetPathValue("workspaceId", "ws-1") + req.SetPathValue("sessionId", "agent-1") + req.Header.Set("Authorization", "Bearer "+signWorkspaceCreateNodeToken(t, key, "node-1", "ws-1")) + req.Header.Set("X-SAM-Node-Id", "node-1") + req.Header.Set("X-SAM-Workspace-Id", "ws-1") + recorder := httptest.NewRecorder() + s.handleHibernateAgentSession(recorder, req) + if recorder.Code != http.StatusAccepted { + t.Fatalf("hibernate status = %d, want 202: %s", recorder.Code, recorder.Body.String()) + } + select { + case <-rec.done: + case <-time.After(45 * time.Second): + t.Fatalf("capture never reported completion or failure; callbacks seen: %v", rec.seen()) + } + return rec.seen() +} + +func TestHibernateDeliveredTokenAuthenticatesEverySnapshotCallback(t *testing.T) { + seen := hibernateWithBody(t, `{"chatSessionId":"chat-1","runtime":"vm","agentType":"openai-codex","background":true,"workspaceCallbackToken":"`+hibernateFreshToken+`"}`) + + for _, callback := range []string{"/session-snapshot/prepare", "/session-snapshot/progress", "/session-snapshot/complete"} { + if len(seen[callback]) == 0 { + t.Fatalf("capture never reached %s; callbacks seen: %v", callback, seen) + } + } + for callback, tokens := range seen { + for _, token := range tokens { + if token != hibernateFreshToken { + t.Fatalf("%s authenticated with %q, want the token delivered with the hibernate request", callback, token) + } + } + } +} + +// Control: without the delivered token the capture uses the expired runtime +// token and fails at prepare, which is the production 401 loop. Proves the test +// above can observe the failure it guards against. +func TestHibernateWithoutDeliveredTokenFailsWithTheExpiredRuntimeToken(t *testing.T) { + seen := hibernateWithBody(t, `{"chatSessionId":"chat-1","runtime":"vm","agentType":"openai-codex","background":true}`) + + prepare := seen["/session-snapshot/prepare"] + if len(prepare) == 0 || prepare[0] != hibernateExpiredToken { + t.Fatalf("prepare tokens = %v, want the expired runtime token", prepare) + } + if len(seen["/session-snapshot/complete"]) != 0 { + t.Fatalf("capture completed without a valid token: %v", seen) + } +} diff --git a/packages/vm-agent/internal/server/workspace_callback_token_renewal_test.go b/packages/vm-agent/internal/server/workspace_callback_token_renewal_test.go new file mode 100644 index 0000000000..7eccaec761 --- /dev/null +++ b/packages/vm-agent/internal/server/workspace_callback_token_renewal_test.go @@ -0,0 +1,533 @@ +package server + +import ( + "database/sql" + "encoding/json" + "net/http" + "net/http/httptest" + "path/filepath" + "strings" + "sync" + "testing" + "time" + + "github.com/golang-jwt/jwt/v5" + + "github.com/workspace/vm-agent/internal/acp" + "github.com/workspace/vm-agent/internal/config" + "github.com/workspace/vm-agent/internal/messagereport" + "github.com/workspace/vm-agent/internal/persistence" +) + +const ( + renewalTestNode = "node-1" + renewalTestNodeToken = "node-token" + renewalTestWorkspace = "ws-1" +) + +// workspaceTestToken mints a token with the production claim set. The agent never +// verifies signatures (the control plane does), so a test key is enough. +func workspaceTestToken(t *testing.T, workspaceID string, issuedAt time.Time, lifetime time.Duration) string { + t.Helper() + token, err := jwt.NewWithClaims(jwt.SigningMethodHS256, jwt.MapClaims{ + "workspace": workspaceID, + "type": "callback", + "scope": "workspace", + "sub": workspaceID, + "aud": "workspace-callback", + "iat": issuedAt.Unix(), + "exp": issuedAt.Add(lifetime).Unix(), + }).SignedString([]byte("test-signing-key")) + if err != nil { + t.Fatal(err) + } + return token +} + +type renewalRequest struct { + workspaceID string + workspaceToken string + nodeID string + nodeToken string +} + +// renewalControlPlane plays the control plane's side of the renewal contract +// (apps/api/src/routes/workspaces/callback-token-renewal.ts): the workspace token +// in Authorization, the node id and node token in the JSON body. It also serves +// the messages endpoint, accepting only tokens it has issued or been told about. +type renewalControlPlane struct { + t *testing.T + mu sync.Mutex + requests []renewalRequest + respond func(req renewalRequest) (int, string) + accepted map[string]bool + messages []string // bearer token of each accepted message batch + rejectedMessages map[string]int + hold chan struct{} + inFlight chan struct{} + holdOnce sync.Once + server *httptest.Server + nodeToken string +} + +func newRenewalControlPlane(t *testing.T) *renewalControlPlane { + cp := &renewalControlPlane{t: t, accepted: map[string]bool{}, rejectedMessages: map[string]int{}, nodeToken: renewalTestNodeToken} + cp.server = httptest.NewServer(http.HandlerFunc(cp.serve)) + t.Cleanup(cp.server.Close) + return cp +} + +func (cp *renewalControlPlane) serve(w http.ResponseWriter, r *http.Request) { + token := strings.TrimPrefix(r.Header.Get("Authorization"), "Bearer ") + switch { + case r.Method == http.MethodPost && strings.HasSuffix(r.URL.Path, "/callback-token/renew"): + var body struct { + NodeID string `json:"nodeId"` + NodeToken string `json:"nodeToken"` + } + _ = json.NewDecoder(r.Body).Decode(&body) + workspaceID := strings.TrimSuffix(strings.TrimPrefix(r.URL.Path, "/api/workspaces/"), "/callback-token/renew") + req := renewalRequest{workspaceID: workspaceID, workspaceToken: token, nodeID: body.NodeID, nodeToken: body.NodeToken} + cp.mu.Lock() + cp.requests = append(cp.requests, req) + hold, inFlight, respond := cp.hold, cp.inFlight, cp.respond + cp.mu.Unlock() + if hold != nil { + cp.holdOnce.Do(func() { close(inFlight) }) + <-hold + } + if body.NodeID != renewalTestNode || body.NodeToken != cp.nodeToken { + writeRenewalJSON(w, http.StatusUnauthorized, `{"error":"NODE_CALLBACK_UNAUTHORIZED","message":"Invalid or expired node callback token"}`) + return + } + status, payload := respond(req) + writeRenewalJSON(w, status, payload) + case r.Method == http.MethodPost && strings.HasSuffix(r.URL.Path, "/messages"): + cp.mu.Lock() + ok := cp.accepted[token] + if ok { + cp.messages = append(cp.messages, token) + } else { + cp.rejectedMessages[token]++ + } + cp.mu.Unlock() + if !ok { + writeRenewalJSON(w, http.StatusUnauthorized, `{"error":"UNAUTHORIZED","message":"Invalid or expired callback token"}`) + return + } + writeRenewalJSON(w, http.StatusOK, `{"persisted":1,"duplicates":0}`) + default: + w.WriteHeader(http.StatusNoContent) + } +} + +func writeRenewalJSON(w http.ResponseWriter, status int, body string) { + w.Header().Set("Content-Type", "application/json") + w.WriteHeader(status) + _, _ = w.Write([]byte(body)) +} + +// renewWith makes the next renewal of token's own workspace return token. Like the +// real route, a renewal is only ever issued for the workspace the request names; +// other workspaces get "not due". +func (cp *renewalControlPlane) renewWith(token string) { + claims := jwt.MapClaims{} + if _, _, err := jwt.NewParser().ParseUnverified(token, claims); err != nil { + cp.t.Fatal(err) + } + workspaceID, _ := claims["workspace"].(string) + cp.mu.Lock() + defer cp.mu.Unlock() + cp.accepted[token] = true + cp.respond = func(req renewalRequest) (int, string) { + if req.workspaceID != workspaceID { + return http.StatusOK, `{"renewed":false}` + } + return http.StatusOK, `{"renewed":true,"token":"` + token + `","expiresAt":"2026-10-06T00:00:00.000Z"}` + } +} + +func (cp *renewalControlPlane) respondWith(status int, body string) { + cp.mu.Lock() + defer cp.mu.Unlock() + cp.respond = func(renewalRequest) (int, string) { return status, body } +} + +// heldOnToken reports whether the messages endpoint has rejected token at least +// once (the reporter then holds its rows until the token changes). +func heldOnToken(cp *renewalControlPlane, token string) bool { + cp.mu.Lock() + defer cp.mu.Unlock() + return cp.rejectedMessages[token] > 0 +} + +func (cp *renewalControlPlane) renewalRequests() []renewalRequest { + cp.mu.Lock() + defer cp.mu.Unlock() + return append([]renewalRequest(nil), cp.requests...) +} + +type renewalHarness struct { + t *testing.T + s *Server + cp *renewalControlPlane + clock time.Time +} + +func newRenewalHarness(t *testing.T, start time.Time) *renewalHarness { + t.Helper() + cp := newRenewalControlPlane(t) + h := &renewalHarness{t: t, cp: cp, clock: start} + h.s = &Server{ + config: &config.Config{ + NodeID: renewalTestNode, + ControlPlaneURL: cp.server.URL, + CallbackToken: renewalTestNodeToken, + HTTPCallbackTimeout: 5 * time.Second, + WorkspaceCallbackTokenRefreshRatio: config.DefaultWorkspaceCallbackTokenRefreshRatio, + WorkspaceCallbackTokenRenewalTimeout: 5 * time.Second, + WorkspaceCallbackTokenRenewalRetryInitial: time.Minute, + WorkspaceCallbackTokenRenewalRetryMax: 30 * time.Minute, + }, + callbackToken: renewalTestNodeToken, + errorReporter: newTestErrorReporter(), + messageReporters: map[string]*messagereport.Reporter{}, + sessionHosts: map[string]*acp.SessionHost{}, + workspaces: map[string]*WorkspaceRuntime{}, + done: make(chan struct{}), + } + h.s.tokenRenewal.now = func() time.Time { return h.clock } + return h +} + +func (h *renewalHarness) addWorkspace(id, token string) *WorkspaceRuntime { + runtime := &WorkspaceRuntime{ID: id, CallbackToken: token, WorkspaceDir: "/workspace/" + id, Status: "running"} + h.s.workspaces[id] = runtime + return runtime +} + +func (h *renewalHarness) token(id string) string { + h.s.workspaceMu.RLock() + defer h.s.workspaceMu.RUnlock() + return h.s.workspaces[id].CallbackToken +} + +// pass advances the clock and runs one post-heartbeat renewal pass. +func (h *renewalHarness) pass(advance time.Duration) { + h.clock = h.clock.Add(advance) + h.s.renewDueWorkspaceCallbackTokensOnce() +} + +var renewalEpoch = time.Date(2026, 10, 2, 20, 20, 0, 0, time.UTC) + +func TestWorkspaceTokenRenewal_RenewsAtTheRefreshPointAndKeepsAWorkspaceAlivePast24h(t *testing.T) { + h := newRenewalHarness(t, renewalEpoch) + first := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour) + h.addWorkspace(renewalTestWorkspace, first) + + h.pass(11 * time.Hour) // 11h of 24h: not due + if got := len(h.cp.renewalRequests()); got != 0 { + t.Fatalf("renewed before the refresh point: %d requests", got) + } + + second := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch.Add(12*time.Hour), 24*time.Hour) + h.cp.renewWith(second) + h.pass(time.Hour) // 12h: due + if h.token(renewalTestWorkspace) != second { + t.Fatal("the renewed token was not installed") + } + + h.pass(11 * time.Hour) // 23h: the renewed token is only 11h old + if got := len(h.cp.renewalRequests()); got != 1 { + t.Fatalf("renewed a fresh token: %d requests", got) + } + + third := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch.Add(24*time.Hour), 24*time.Hour) + h.cp.renewWith(third) + h.pass(time.Hour) // 24h: the original token has expired; the renewed one is due + h.pass(time.Hour) // 25h + if h.token(renewalTestWorkspace) != third { + t.Fatal("the workspace does not hold a live token past its first 24h") + } + + requests := h.cp.renewalRequests() + if len(requests) != 2 || requests[0].workspaceToken != first || requests[1].workspaceToken != second { + t.Fatalf("renewals presented %v, want the first then the second token", requests) + } + for _, req := range requests { + if req.workspaceID != renewalTestWorkspace || req.nodeID != renewalTestNode || req.nodeToken != renewalTestNodeToken { + t.Fatalf("renewal did not carry the workspace and node proofs: %+v", req) + } + } + if h.s.getCallbackToken() != renewalTestNodeToken { + t.Fatal("workspace renewal must not touch the node token") + } +} + +func TestWorkspaceTokenRenewal_RefusalLatchesUntilANewTokenArrives(t *testing.T) { + for _, tc := range []struct { + name string + status int + body string + }{ + {"expired workspace token", http.StatusUnauthorized, `{"error":"UNAUTHORIZED","message":"Invalid or expired callback token"}`}, + {"not hosted on this node", http.StatusForbidden, `{"error":"FORBIDDEN","message":"Workspace is not hosted on this node"}`}, + {"workspace deleted", http.StatusGone, `{"error":"GONE","message":"Workspace is deleted; callback resource is gone"}`}, + } { + t.Run(tc.name, func(t *testing.T) { + h := newRenewalHarness(t, renewalEpoch) + first := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour) + h.addWorkspace(renewalTestWorkspace, first) + h.cp.respondWith(tc.status, tc.body) + + h.pass(13 * time.Hour) + h.pass(time.Hour) + h.pass(48 * time.Hour) + if got := len(h.cp.renewalRequests()); got != 1 { + t.Fatalf("a refused token was presented %d times, want once", got) + } + if h.token(renewalTestWorkspace) != first { + t.Fatal("a refusal must not change the workspace token") + } + + // A control-plane delivery (e.g. hibernate) brings a new token: renewal + // resumes for it, from its own refresh point. + delivered := workspaceTestToken(t, renewalTestWorkspace, h.clock, 24*time.Hour) + h.s.upsertWorkspaceRuntime(renewalTestWorkspace, "", "", "", delivered) + h.cp.renewWith(workspaceTestToken(t, renewalTestWorkspace, h.clock.Add(12*time.Hour), 24*time.Hour)) + h.pass(time.Hour) + if got := len(h.cp.renewalRequests()); got != 1 { + t.Fatalf("a fresh delivered token was renewed early: %d requests", got) + } + h.pass(11 * time.Hour) + requests := h.cp.renewalRequests() + if len(requests) != 2 || requests[1].workspaceToken != delivered { + t.Fatalf("renewal did not resume with the delivered token: %v", requests) + } + }) + } +} + +func TestWorkspaceTokenRenewal_TransientFailuresBackOffAndNodeCredentialFailuresRetry(t *testing.T) { + h := newRenewalHarness(t, renewalEpoch) + h.addWorkspace(renewalTestWorkspace, workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour)) + h.cp.respondWith(http.StatusServiceUnavailable, `{"error":"SERVICE_UNAVAILABLE"}`) + + h.pass(13 * time.Hour) // attempt 1 fails → wait 1m + h.pass(30 * time.Second) + h.pass(30 * time.Second) // attempt 2 fails → wait 2m + h.pass(time.Minute) + h.pass(time.Minute) // attempt 3 fails → wait 4m + if got := len(h.cp.renewalRequests()); got != 3 { + t.Fatalf("transient failures: %d attempts, want 3 on a 1m/2m backoff", got) + } + h.pass(10 * time.Hour) // the backoff is capped at RetryMax + if got := len(h.cp.renewalRequests()); got != 4 { + t.Fatalf("attempts after a long wait = %d, want 4", got) + } + + // A refused NODE credential is the node token's problem (the heartbeat + // refreshes it), so it is retried rather than latched. + h.cp.mu.Lock() + h.cp.nodeToken = "a-node-token-the-agent-does-not-have-yet" + h.cp.mu.Unlock() + h.pass(31 * time.Minute) + h.cp.mu.Lock() + h.cp.nodeToken = renewalTestNodeToken + h.cp.mu.Unlock() + renewed := workspaceTestToken(t, renewalTestWorkspace, h.clock, 24*time.Hour) + h.cp.renewWith(renewed) + h.pass(31 * time.Minute) + if h.token(renewalTestWorkspace) != renewed { + t.Fatal("renewal did not recover after the node credential was accepted again") + } +} + +func TestWorkspaceTokenRenewal_NotDueWaitsRetryMax(t *testing.T) { + h := newRenewalHarness(t, renewalEpoch) + h.addWorkspace(renewalTestWorkspace, workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour)) + h.cp.respondWith(http.StatusOK, `{"renewed":false}`) // e.g. the agent clock runs ahead + + h.pass(13 * time.Hour) + h.pass(29 * time.Minute) + if got := len(h.cp.renewalRequests()); got != 1 { + t.Fatalf("asked again before RetryMax: %d requests", got) + } + h.pass(time.Minute) + if got := len(h.cp.renewalRequests()); got != 2 { + t.Fatalf("did not ask again after RetryMax: %d requests", got) + } +} + +func TestWorkspaceTokenRenewal_ExpiredTokenIsOfferedOnceSoTheControlPlaneDecides(t *testing.T) { + h := newRenewalHarness(t, renewalEpoch) + h.addWorkspace(renewalTestWorkspace, workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour)) + h.cp.respondWith(http.StatusUnauthorized, `{"error":"UNAUTHORIZED"}`) + + h.pass(30 * time.Hour) // the agent believes the token expired 6h ago + h.pass(time.Hour) + if got := len(h.cp.renewalRequests()); got != 1 { + t.Fatalf("an apparently expired token was offered %d times, want exactly once", got) + } +} + +func TestWorkspaceTokenRenewal_DiscardsARenewalWhenADeliveryWinsTheRace(t *testing.T) { + h := newRenewalHarness(t, renewalEpoch) + first := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour) + h.addWorkspace(renewalTestWorkspace, first) + host := acp.NewSessionHost(acp.SessionHostConfig{GatewayConfig: acp.GatewayConfig{CallbackToken: first}}) + t.Cleanup(host.Stop) + h.s.sessionHosts[renewalTestWorkspace+":agent-1"] = host + + renewed := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch.Add(13*time.Hour), 24*time.Hour) + h.cp.renewWith(renewed) + h.cp.mu.Lock() + h.cp.hold, h.cp.inFlight = make(chan struct{}), make(chan struct{}) + hold, inFlight := h.cp.hold, h.cp.inFlight + h.cp.mu.Unlock() + + h.clock = h.clock.Add(13 * time.Hour) + done := make(chan struct{}) + go func() { + h.s.renewDueWorkspaceCallbackTokensOnce() + close(done) + }() + <-inFlight + // The control plane delivers a token over hibernate while the renewal is in flight. + delivered := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch.Add(13*time.Hour+time.Second), 24*time.Hour) + h.s.upsertWorkspaceRuntime(renewalTestWorkspace, "", "", "", delivered) + close(hold) + <-done + + if h.token(renewalTestWorkspace) != delivered { + t.Fatal("a late renewal response overwrote a token delivered during the request") + } + if !host.UsesCallbackToken(delivered) { + t.Fatal("the SessionHost did not receive the delivered token") + } +} + +func TestWorkspaceTokenDelivery_NeverAdoptsATokenThatExpiresEarlier(t *testing.T) { + h := newRenewalHarness(t, renewalEpoch) + current := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch.Add(12*time.Hour), 24*time.Hour) + h.addWorkspace(renewalTestWorkspace, current) + host := acp.NewSessionHost(acp.SessionHostConfig{GatewayConfig: acp.GatewayConfig{CallbackToken: current}}) + t.Cleanup(host.Stop) + h.s.sessionHosts[renewalTestWorkspace+":agent-1"] = host + + older := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour) // a reordered, older delivery + h.s.upsertWorkspaceRuntime(renewalTestWorkspace, "", "", "", older) + if h.token(renewalTestWorkspace) != current || !host.UsesCallbackToken(current) { + t.Fatal("an older delivered token replaced a newer one") + } + + newer := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch.Add(20*time.Hour), 24*time.Hour) + h.s.upsertWorkspaceRuntime(renewalTestWorkspace, "", "", "", newer) + if h.token(renewalTestWorkspace) != newer || !host.UsesCallbackToken(newer) { + t.Fatal("a newer delivered token was not adopted and propagated") + } +} + +// Renewal must reach every consumer that copied the old token, and a parked +// message reporter must deliver what it held. +func TestWorkspaceTokenRenewal_ReachesParkedReporterAndEverySessionHostOfTheWorkspace(t *testing.T) { + h := newRenewalHarness(t, renewalEpoch) + first := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour) + h.addWorkspace(renewalTestWorkspace, first) + other := workspaceTestToken(t, "ws-2", renewalEpoch, 24*time.Hour) + h.addWorkspace("ws-2", other) + + hosts := map[string]*acp.SessionHost{} + for _, key := range []string{renewalTestWorkspace + ":agent-1", renewalTestWorkspace + ":agent-2", "ws-2:agent-1"} { + token := first + if strings.HasPrefix(key, "ws-2:") { + token = other + } + host := acp.NewSessionHost(acp.SessionHostConfig{GatewayConfig: acp.GatewayConfig{CallbackToken: token}}) + t.Cleanup(host.Stop) + hosts[key] = host + h.s.sessionHosts[key] = host + } + + _, db := openTestSQLiteDB(t) + reporter, err := messagereport.New(db, messagereport.Config{ + BatchMaxWait: 20 * time.Millisecond, Endpoint: h.cp.server.URL, + WorkspaceID: renewalTestWorkspace, ProjectID: "proj-1", SessionID: "chat-1", + }) + if err != nil { + t.Fatal(err) + } + t.Cleanup(reporter.Shutdown) + reporter.SetToken(first) + h.s.messageReporters[renewalTestWorkspace] = reporter + if err := reporter.Enqueue(messagereport.Message{MessageID: "m-1", Role: "assistant", Content: "reply"}); err != nil { + t.Fatal(err) + } + // The control plane rejects the old token: the reporter holds the message. + waitFor(t, func() bool { return countOutbox(t, db) == 1 && heldOnToken(h.cp, first) }, "message held after 401") + + renewed := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch.Add(12*time.Hour), 24*time.Hour) + h.cp.renewWith(renewed) + h.pass(12 * time.Hour) + + waitFor(t, func() bool { return countOutbox(t, db) == 0 }, "held message delivered after renewal") + h.cp.mu.Lock() + delivered := append([]string(nil), h.cp.messages...) + h.cp.mu.Unlock() + if len(delivered) != 1 || delivered[0] != renewed { + t.Fatalf("held message delivered with %v, want once with the renewed token", delivered) + } + if !hosts[renewalTestWorkspace+":agent-1"].UsesCallbackToken(renewed) || + !hosts[renewalTestWorkspace+":agent-2"].UsesCallbackToken(renewed) { + t.Fatal("a SessionHost of the renewed workspace kept the old token") + } + if !hosts["ws-2:agent-1"].UsesCallbackToken(other) { + t.Fatal("another workspace's SessionHost received this workspace's token") + } +} + +func TestWorkspaceTokenRenewal_PersistsTheRenewedTokenForRestart(t *testing.T) { + path := filepath.Join(t.TempDir(), "state.db") + openStore := func() *persistence.Store { + store, err := persistence.Open(path) + if err != nil { + t.Fatal(err) + } + if err := store.SetCallbackTokenEncryptionSecret(renewalTestNodeToken); err != nil { + t.Fatal(err) + } + return store + } + + h := newRenewalHarness(t, renewalEpoch) + h.s.store = openStore() + first := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour) + runtime := h.addWorkspace(renewalTestWorkspace, first) + runtime.Repository = "octo/repo" + h.s.persistWorkspaceMetadata(runtime) + + renewed := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch.Add(12*time.Hour), 24*time.Hour) + h.cp.renewWith(renewed) + h.pass(12 * time.Hour) + if err := h.s.store.Close(); err != nil { + t.Fatal(err) + } + + // A restarted agent hydrates the workspace from SQLite. + restarted := newRenewalHarness(t, h.clock) + restarted.s.store = openStore() + t.Cleanup(func() { _ = restarted.s.store.Close() }) + restarted.s.upsertWorkspaceRuntime(renewalTestWorkspace, "", "", "", "") + if restarted.token(renewalTestWorkspace) != renewed { + t.Fatal("a restart resumed with the pre-renewal token") + } +} + +func countOutbox(t *testing.T, db *sql.DB) int { + t.Helper() + var count int + if err := db.QueryRow("SELECT COUNT(*) FROM message_outbox").Scan(&count); err != nil { + t.Fatalf("count outbox: %v", err) + } + return count +} From 1fcfaa4b54d2c0ed7372952dc6fcdc7da6402d52 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 09:06:01 +0000 Subject: [PATCH 06/19] docs: callback token scopes, lifetime and workspace renewal Document the two callback token scopes, the 24h lifetime (the security page said minutes), dual-proof workspace renewal and VM-only control-plane delivery, plus the new agent and API configuration. Add a chat-session rebind race case for delivery and record discrimination evidence in the task file. Co-Authored-By: Claude Opus 5.5 --- .claude/skills/api-reference/SKILL.md | 1 + .claude/skills/env-reference/SKILL.md | 15 +++ .../workspace-callback-token-renewal.test.ts | 12 ++ .../docs/docs/architecture/security.md | 28 ++++- .../content/docs/docs/reference/vm-agent.md | 105 +++++++++--------- ...-10-04-workspace-callback-token-renewal.md | 74 +++++++----- 6 files changed, 151 insertions(+), 84 deletions(-) diff --git a/.claude/skills/api-reference/SKILL.md b/.claude/skills/api-reference/SKILL.md index 4753048675..624c8da8b3 100644 --- a/.claude/skills/api-reference/SKILL.md +++ b/.claude/skills/api-reference/SKILL.md @@ -220,6 +220,7 @@ The MCP `create_trigger` tool intentionally creates cron triggers only. Generic - `POST /api/workspaces/:id/provisioning-failed` — Workspace provisioning failure callback (sets workspace to `error`) - `POST /api/workspaces/:id/heartbeat` — Workspace activity heartbeat callback - `GET /api/workspaces/:id/runtime` — Workspace runtime metadata callback (repository/branch for recovery) +- `POST /api/workspaces/:id/callback-token/renew` — Renew a workspace callback token before it expires. Two proofs: the current, unexpired workspace token in `Authorization`, and `{ nodeId, nodeToken }` (the hosting node's node-scoped token) in the body. Renews only while the workspace is `creating`/`running`/`recovery`, bound to that node, owned by the node's user, on a non-terminal node, and past `CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO` of the token's lifetime (otherwise `{ renewed: false }`). The renewed token keeps the chain's first `iat` as `gen_iat`. Returns `{ renewed: true, token, expiresAt }` with `Cache-Control: no-store`; 401/403/410 otherwise (`NODE_CALLBACK_UNAUTHORIZED`/`NODE_CALLBACK_FORBIDDEN` when only the node proof failed) - `POST /api/workspaces/:id/boot-log` — Workspace boot progress log callback - `POST /api/workspaces/:id/agent-settings` — Workspace agent settings callback (model, permissionMode) - `POST /api/projects/:id/workspaces/:workspaceId/eviction` — VM-agent callback JWT endpoint that validates node/workspace/runtime-generation identity and successful container stop, atomically marks the workspace `evicted` and closes usage/agent sessions, then serializes replay-safe ProjectData finalization through NodeLifecycle. A stale generation returns 410; failed finalization is retryable. Explicit restart requires renewed capacity admission and rotates the runtime generation diff --git a/.claude/skills/env-reference/SKILL.md b/.claude/skills/env-reference/SKILL.md index b83f50d529..6d44dcfb72 100644 --- a/.claude/skills/env-reference/SKILL.md +++ b/.claude/skills/env-reference/SKILL.md @@ -534,6 +534,11 @@ by the read-only cron-liveness check. - `MCP_INCIDENT_LIST_MAX` — Maximum result count accepted by the private `list_incident_queue` MCP tool (default: 50) - `HETZNER_MAX_LIST_PAGES` — Maximum pages per Hetzner list request (default: 100) +### Callback Tokens + +- `CALLBACK_TOKEN_EXPIRY_MS` — Lifetime of node- and workspace-scoped VM callback JWTs (default: `86400000` / 24h). Changing it does not extend tokens already issued. +- `CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO` — Fraction of a callback token's lifetime after which it may be renewed (default: `0.5`, clamped to `0.1`–`0.9`). Gates both the node token refresh in `POST /api/nodes/:id/heartbeat` and workspace token renewal in `POST /api/workspaces/:id/callback-token/renew`; a token younger than this is not re-minted. + ### Timeouts - `ORCHESTRATOR_STOP_CAS_MAX_ATTEMPTS` — Maximum task-status compare-and-set attempts after a parent hard-stops a child runtime (default: 2) @@ -727,6 +732,16 @@ Generated deployments validate and pass these values through cloud-init to newly - `SESSION_SNAPSHOT_PROGRESS_REPORT_INTERVAL` — Minimum interval between progress callbacks while a checkpoint continues making progress (default: `15s`) - `SESSION_SNAPSHOT_PROGRESS_REPORT_TIMEOUT` — Timeout for each best-effort progress callback to the control plane (default: `5s`) +### Workspace Callback Token Renewal + +The agent renews each workspace callback token after a successful node heartbeat once the token is past the refresh ratio (`internal/server/workspace_callback_token_renewal.go`). These use their defaults unless set in the agent service environment. + +- `WORKSPACE_CALLBACK_TOKEN_REFRESH_RATIO` — Fraction of a workspace token's lifetime after which the agent renews it (default: `0.5`, clamped to `0.1`–`0.9`; the control plane's `CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO` still decides) +- `WORKSPACE_CALLBACK_TOKEN_RENEWAL_TIMEOUT` — Timeout for one renewal request (default: `15s`) +- `WORKSPACE_CALLBACK_TOKEN_RENEWAL_RETRY_INITIAL` — First backoff after a transient renewal failure (default: `1m`) +- `WORKSPACE_CALLBACK_TOKEN_RENEWAL_RETRY_MAX` — Backoff ceiling, also the wait after a "not yet due" answer (default: `30m`) +- `MSG_AUTH_RENEWAL_WAIT` — How long chat-message delivery may stay paused on a rejected (401) workspace token before the pause is reported as an error (default: `15m`). Held messages are kept either way. + ### File Operations - `FILE_LIST_TIMEOUT` — Timeout for file listing commands (default: 10s) diff --git a/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts b/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts index e7db3ad5af..9f93b7ded4 100644 --- a/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts +++ b/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts @@ -148,6 +148,18 @@ describe('mintWorkspaceCallbackTokenForNodeDelivery', () => { ).resolves.toBeNull(); }); + it('withholds the token when the workspace is rebound to another chat session after the first read', async () => { + const reads = bindingReadCount(); + reads.mutateOn = (n) => { + if (n === 2) { + sqlite.prepare("UPDATE workspaces SET chat_session_id = 'chat-2' WHERE id = ?").run(WS); + } + }; + await expect( + mintWorkspaceCallbackTokenForNodeDelivery(makeEnv(), { workspaceId: WS, nodeId: NODE }) + ).resolves.toBeNull(); + }); + it('never delivers to an Instant (cf-container) runtime', async () => { sqlite.prepare("UPDATE nodes SET runtime = 'cf-container' WHERE id = ?").run(NODE); await expect( diff --git a/apps/www/src/content/docs/docs/architecture/security.md b/apps/www/src/content/docs/docs/architecture/security.md index cab7a1ff6b..70eaee272e 100644 --- a/apps/www/src/content/docs/docs/architecture/security.md +++ b/apps/www/src/content/docs/docs/architecture/security.md @@ -91,12 +91,28 @@ SAM uses **BetterAuth** with configured OAuth login providers for user authentic ### Token Types -| Token | Lifetime | Purpose | Validated By | -| --------------- | --------- | -------------------------------- | ----------------------- | -| Session cookie | Hours | Browser authentication | API Worker (BetterAuth) | -| Workspace JWT | Minutes | Terminal WebSocket auth | VM Agent (via JWKS) | -| Bootstrap token | 5 minutes | One-time VM credential injection | API Worker | -| Callback token | Minutes | VM Agent → API callbacks | API Worker | +| Token | Lifetime | Purpose | Validated By | +| --------------- | -------------------- | -------------------------------- | ----------------------- | +| Session cookie | Hours | Browser authentication | API Worker (BetterAuth) | +| Workspace JWT | Minutes | Terminal WebSocket auth | VM Agent (via JWKS) | +| Bootstrap token | 5 minutes | One-time VM credential injection | API Worker | +| Callback token | 24 hours (renewable) | VM Agent → API callbacks | API Worker | + +### Callback Tokens + +VM agents call the API Worker with RS256 callback tokens signed by the Worker (`apps/api/src/services/jwt.ts`). There are two scopes, and neither can stand in for the other: + +- **Node tokens** (`scope: node`) authenticate node-level callbacks such as heartbeats and error reports. They are renewed in the heartbeat response while the node is not terminal (`apps/api/src/routes/node-lifecycle.ts`). +- **Workspace tokens** (`scope: workspace`) authenticate everything about one workspace: chat messages, session snapshots, git credentials, runtime assets, task status and ACP activity. The Worker hands one to the node when it creates or restores the workspace. + +Both last `CALLBACK_TOKEN_EXPIRY_MS` (24 hours by default). A workspace token is renewed without ever leaving its workspace's scope: + +1. **Renewal.** After each successful heartbeat, the VM agent renews workspace tokens that are past `CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO` of their lifetime by calling `POST /api/workspaces/:id/callback-token/renew` with two proofs: the workspace's current, unexpired token and the node's own token. The Worker renews only when D1 binds the workspace to that node, the node belongs to the workspace's owner, and both are still active (`apps/api/src/services/workspace-callback-token-renewal.ts`). A node token alone cannot obtain a workspace token, a workspace token copied out of a devcontainer cannot renew itself, and an expired token is never renewed. +2. **Delivery.** When the Worker asks a VM node to snapshot a session for sleep, the request carries a fresh workspace token over the authenticated node-management channel, the same way workspace creation does. It is minted only if the workspace is still active on that node, and never for Instant containers, which receive a fresh token on every cold wake. + +Deleting, stopping or moving a workspace ends renewal, so its callback authority still lapses within one token lifetime. A renewed token keeps its first issue time in a `gen_iat` claim, so the Instant stale-callback guard still recognizes a callback from a replaced container (`apps/api/src/routes/_stale-callback-guard.ts`). + +Credentials that an agent process received when it started, such as the SAM AI proxy key, are not rotated inside the running process; the process picks up the current token the next time it starts. Deletion-in-progress callbacks fail closed. A VM delete timeout is treated as uncertainty, so the workspace remains `stopping` and callback routes reject its diff --git a/apps/www/src/content/docs/docs/reference/vm-agent.md b/apps/www/src/content/docs/docs/reference/vm-agent.md index 959006e28e..2300525886 100644 --- a/apps/www/src/content/docs/docs/reference/vm-agent.md +++ b/apps/www/src/content/docs/docs/reference/vm-agent.md @@ -256,56 +256,61 @@ Validates workspace and node-management JWTs using the API's JWKS endpoint: The agent reads the following environment variables. Cloud-init supplies node identity and callback configuration; monitoring values use these defaults unless overridden in the agent service environment: -| Variable | Default | Description | -| ---------------------------------------------- | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `NODE_ID` | — | Unique node identifier | -| `CONTROL_PLANE_URL` | — | API Worker URL for callbacks | -| `CALLBACK_TOKEN_FILE` | `/etc/sam/callback-token` on cloud-init nodes | Root-only file containing the callback JWT for authenticating callbacks. `CALLBACK_TOKEN` remains a legacy fallback for already-provisioned nodes/manual runs. | -| `LOG_LEVEL` | `info` | Log level: `debug`, `info`, `warn`, `error` | -| `LOG_FORMAT` | `json` | Output format: `json` or `text` | -| `ACP_PROMPT_RETRY_MAX_RETRIES` | `2` | Max transient provider prompt retries after the initial attempt | -| `ACP_PROMPT_RETRY_INITIAL_BACKOFF` | `15s` | Initial backoff before retrying transient provider prompt errors | -| `ACP_PROMPT_RETRY_MAX_BACKOFF` | `2m` | Max exponential backoff for transient provider prompt retries | -| `ACP_CHECKPOINT_PREEMPT_GRACE` | `30s` | Grace after ACP cancel/close before force-stopping the harness | -| `ACP_CHECKPOINT_PREEMPT_MAX_GRACE` | `2m` | Maximum `graceMs` accepted by the rollover endpoint | -| `ACP_CHECKPOINT_ROLLOVER_TIMEOUT` | `2m` | Deadline for the complete stop, restart, and strict LoadSession operation | -| `ACP_NOTIF_SERIALIZE_TIMEOUT` | `5s` | Timeout for ACP notification serialization | -| `ACP_HARNESS_ACTIVITY_REPORT_DEBOUNCE` | `750ms` | Debounce window for coalescing ACP harness/tool-call activity reports before POSTing activity callbacks | -| `STANDALONE_CLONE_FILTER` | `blob:none` | Git partial-clone filter for standalone (Cloudflare Container) workspace clones, which run synchronously inside the control plane's create-workspace request (`cloneStandaloneRepository` in `internal/server/standalone_workspace.go`). Set `off` to force full clones. The control plane forwards `CF_CONTAINER_CLONE_FILTER` here. | -| `GRACEFUL_SHUTDOWN_TIMEOUT` | `30s` | Max time to wait for VM-agent HTTP server shutdown after SIGTERM | -| `SYSTEM_PROVISIONING_TIMEOUT` | `15m` | Max time for workspace host provisioning before bootstrap | -| `CF_IP_FETCH_TIMEOUT` | `10s` | Timeout for fetching Cloudflare IP ranges during firewall provisioning | -| `BOOT_LOG_HTTP_TIMEOUT` | `10s` | Timeout for boot-log callbacks to the control plane | -| `MCP_SHORT_COMMAND_TIMEOUT` | `10s` | Timeout for short MCP workspace probes such as branch and credential checks | -| `MCP_DIFF_COMMAND_TIMEOUT` | `30s` | Timeout for MCP diff-summary git commands | -| `MCP_BUILD_PREPARE_TIMEOUT` | `30s` | Timeout for MCP build/publish preparation probes | -| `JWKS_FETCH_TIMEOUT` | `10s` | Timeout for VM-agent startup JWKS fetches | -| `ACP_CREDENTIAL_SYNC_TIMEOUT` | `10s` | Timeout for ACP auth-file sync-back during shutdown | -| `ACP_RESTART_ATTEMPT_TIMEOUT` | `5m` | Bounds one automatic agent restart attempt by the ACP process monitor | -| `ACP_ACTIVITY_REPORT_TIMEOUT` | `10s` | Timeout for each ACP activity callback attempt | -| `DEVCONTAINER_CACHE_PUSH_TIMEOUT` | `10m` | Timeout for best-effort devcontainer cache image pushes | -| `WORKSPACE_BUILD_QUEUE_DEPTH` | `1` | Concurrent devcontainer build slots on each workspace VM. Supported values are `1` through `16`; invalid values do not enable additional build slots and fall back to the default one-slot behavior. New cloud-init nodes receive this from the Worker env; older agents that do not know it keep their built-in one-slot queue. | -| `DEPLOY_PREFLIGHT_COMMAND_TIMEOUT` | `15s` | Timeout for deployment preflight diagnostic commands | -| `LOG_STREAM_PING_WRITE_TIMEOUT` | `10s` | Write deadline for log-stream WebSocket ping frames | -| `DEFAULT_RESOURCE_EVENT_BUFFER_SIZE` | `64` | Capacity of each bounded pressure/Docker event queue; positive integer | -| `DEFAULT_PSI_POLL_INTERVAL_SECONDS` | `10` | Linux memory PSI sampling interval, in seconds | -| `DEFAULT_CONTAINER_STATS_INTERVAL_SECONDS` | `30` | Docker resource statistics sampling interval, in seconds | -| `DEFAULT_PSI_MEMORY_SOME_WARNING_THRESHOLD` | `25` | Warning threshold for the maximum PSI some-memory avg10/avg60 percentage | -| `DEFAULT_PSI_MEMORY_SOME_CRITICAL_THRESHOLD` | `50` | Critical threshold for the maximum PSI some-memory avg10/avg60 percentage | -| `DEFAULT_PSI_MEMORY_FULL_WARNING_THRESHOLD` | `10` | Warning threshold for the maximum PSI full-memory avg10/avg60 percentage | -| `DEFAULT_PSI_MEMORY_FULL_CRITICAL_THRESHOLD` | `25` | Critical threshold for the maximum PSI full-memory avg10/avg60 percentage | -| `DEFAULT_EVICTION_DEBOUNCE_SECONDS` | `30` | Minimum cooldown between ResourceGuard eviction attempts, in seconds | -| `DEFAULT_EVICTION_SNAPSHOT_TIMEOUT_SECONDS` | `120` | Deadline for pre-stop ResourceGuard eviction snapshot capture, in seconds | -| `DEFAULT_EVICTION_DOCKER_STOP_TIMEOUT_SECONDS` | `10` | Grace period passed to `docker stop --time` during ResourceGuard eviction, in seconds | -| `DEFAULT_EVICTION_CALLBACK_RETRY_MAX_SECONDS` | `300` | Backoff cap for durable eviction callback retries, in seconds; the operation lease is a lower bound and can exceed this cap. Delivery starts on a later heartbeat | -| `DEFAULT_EVICTION_RESOLVE_TIMEOUT_SECONDS` | `5` | Deadline for resolving a pressured Docker container to a workspace before eviction, in seconds | -| `COMPOSE_OUTPUT_RETENTION_BYTES` | `65536` | Retained tail of a deployment compose command's combined output, in bytes (valid range 1024–1048576) | -| `RESOURCE_HISTORY_SAMPLE_INTERVAL` | `5s` | Retained resource-history cgroup sampling cadence | -| `RESOURCE_HISTORY_CHUNK_INTERVAL` | `15m` | Retained resource-history chunk duration before upload | -| `RESOURCE_HISTORY_SPOOL_DIR` | `/var/lib/vm-agent/resource-history` | Node-local retry spool for resource-history chunks | -| `RESOURCE_HISTORY_SPOOL_MAX_BYTES` | `20971520` | Max node-local resource-history retry spool bytes | -| `RESOURCE_HISTORY_UPLOAD_TIMEOUT` | `10s` | Deadline for one resource-history upload callback | -| `RESOURCE_HISTORY_MAX_SAMPLES` | `4096` | Max resource samples packed into one uploaded chunk | +| Variable | Default | Description | +| ------------------------------------------------ | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `NODE_ID` | — | Unique node identifier | +| `CONTROL_PLANE_URL` | — | API Worker URL for callbacks | +| `CALLBACK_TOKEN_FILE` | `/etc/sam/callback-token` on cloud-init nodes | Root-only file containing the callback JWT for authenticating callbacks. `CALLBACK_TOKEN` remains a legacy fallback for already-provisioned nodes/manual runs. | +| `LOG_LEVEL` | `info` | Log level: `debug`, `info`, `warn`, `error` | +| `LOG_FORMAT` | `json` | Output format: `json` or `text` | +| `ACP_PROMPT_RETRY_MAX_RETRIES` | `2` | Max transient provider prompt retries after the initial attempt | +| `ACP_PROMPT_RETRY_INITIAL_BACKOFF` | `15s` | Initial backoff before retrying transient provider prompt errors | +| `ACP_PROMPT_RETRY_MAX_BACKOFF` | `2m` | Max exponential backoff for transient provider prompt retries | +| `ACP_CHECKPOINT_PREEMPT_GRACE` | `30s` | Grace after ACP cancel/close before force-stopping the harness | +| `ACP_CHECKPOINT_PREEMPT_MAX_GRACE` | `2m` | Maximum `graceMs` accepted by the rollover endpoint | +| `ACP_CHECKPOINT_ROLLOVER_TIMEOUT` | `2m` | Deadline for the complete stop, restart, and strict LoadSession operation | +| `ACP_NOTIF_SERIALIZE_TIMEOUT` | `5s` | Timeout for ACP notification serialization | +| `ACP_HARNESS_ACTIVITY_REPORT_DEBOUNCE` | `750ms` | Debounce window for coalescing ACP harness/tool-call activity reports before POSTing activity callbacks | +| `STANDALONE_CLONE_FILTER` | `blob:none` | Git partial-clone filter for standalone (Cloudflare Container) workspace clones, which run synchronously inside the control plane's create-workspace request (`cloneStandaloneRepository` in `internal/server/standalone_workspace.go`). Set `off` to force full clones. The control plane forwards `CF_CONTAINER_CLONE_FILTER` here. | +| `GRACEFUL_SHUTDOWN_TIMEOUT` | `30s` | Max time to wait for VM-agent HTTP server shutdown after SIGTERM | +| `SYSTEM_PROVISIONING_TIMEOUT` | `15m` | Max time for workspace host provisioning before bootstrap | +| `CF_IP_FETCH_TIMEOUT` | `10s` | Timeout for fetching Cloudflare IP ranges during firewall provisioning | +| `BOOT_LOG_HTTP_TIMEOUT` | `10s` | Timeout for boot-log callbacks to the control plane | +| `MCP_SHORT_COMMAND_TIMEOUT` | `10s` | Timeout for short MCP workspace probes such as branch and credential checks | +| `MCP_DIFF_COMMAND_TIMEOUT` | `30s` | Timeout for MCP diff-summary git commands | +| `MCP_BUILD_PREPARE_TIMEOUT` | `30s` | Timeout for MCP build/publish preparation probes | +| `JWKS_FETCH_TIMEOUT` | `10s` | Timeout for VM-agent startup JWKS fetches | +| `ACP_CREDENTIAL_SYNC_TIMEOUT` | `10s` | Timeout for ACP auth-file sync-back during shutdown | +| `ACP_RESTART_ATTEMPT_TIMEOUT` | `5m` | Bounds one automatic agent restart attempt by the ACP process monitor | +| `ACP_ACTIVITY_REPORT_TIMEOUT` | `10s` | Timeout for each ACP activity callback attempt | +| `DEVCONTAINER_CACHE_PUSH_TIMEOUT` | `10m` | Timeout for best-effort devcontainer cache image pushes | +| `WORKSPACE_BUILD_QUEUE_DEPTH` | `1` | Concurrent devcontainer build slots on each workspace VM. Supported values are `1` through `16`; invalid values do not enable additional build slots and fall back to the default one-slot behavior. New cloud-init nodes receive this from the Worker env; older agents that do not know it keep their built-in one-slot queue. | +| `DEPLOY_PREFLIGHT_COMMAND_TIMEOUT` | `15s` | Timeout for deployment preflight diagnostic commands | +| `LOG_STREAM_PING_WRITE_TIMEOUT` | `10s` | Write deadline for log-stream WebSocket ping frames | +| `DEFAULT_RESOURCE_EVENT_BUFFER_SIZE` | `64` | Capacity of each bounded pressure/Docker event queue; positive integer | +| `DEFAULT_PSI_POLL_INTERVAL_SECONDS` | `10` | Linux memory PSI sampling interval, in seconds | +| `DEFAULT_CONTAINER_STATS_INTERVAL_SECONDS` | `30` | Docker resource statistics sampling interval, in seconds | +| `DEFAULT_PSI_MEMORY_SOME_WARNING_THRESHOLD` | `25` | Warning threshold for the maximum PSI some-memory avg10/avg60 percentage | +| `DEFAULT_PSI_MEMORY_SOME_CRITICAL_THRESHOLD` | `50` | Critical threshold for the maximum PSI some-memory avg10/avg60 percentage | +| `DEFAULT_PSI_MEMORY_FULL_WARNING_THRESHOLD` | `10` | Warning threshold for the maximum PSI full-memory avg10/avg60 percentage | +| `DEFAULT_PSI_MEMORY_FULL_CRITICAL_THRESHOLD` | `25` | Critical threshold for the maximum PSI full-memory avg10/avg60 percentage | +| `DEFAULT_EVICTION_DEBOUNCE_SECONDS` | `30` | Minimum cooldown between ResourceGuard eviction attempts, in seconds | +| `DEFAULT_EVICTION_SNAPSHOT_TIMEOUT_SECONDS` | `120` | Deadline for pre-stop ResourceGuard eviction snapshot capture, in seconds | +| `DEFAULT_EVICTION_DOCKER_STOP_TIMEOUT_SECONDS` | `10` | Grace period passed to `docker stop --time` during ResourceGuard eviction, in seconds | +| `DEFAULT_EVICTION_CALLBACK_RETRY_MAX_SECONDS` | `300` | Backoff cap for durable eviction callback retries, in seconds; the operation lease is a lower bound and can exceed this cap. Delivery starts on a later heartbeat | +| `DEFAULT_EVICTION_RESOLVE_TIMEOUT_SECONDS` | `5` | Deadline for resolving a pressured Docker container to a workspace before eviction, in seconds | +| `COMPOSE_OUTPUT_RETENTION_BYTES` | `65536` | Retained tail of a deployment compose command's combined output, in bytes (valid range 1024–1048576) | +| `RESOURCE_HISTORY_SAMPLE_INTERVAL` | `5s` | Retained resource-history cgroup sampling cadence | +| `RESOURCE_HISTORY_CHUNK_INTERVAL` | `15m` | Retained resource-history chunk duration before upload | +| `RESOURCE_HISTORY_SPOOL_DIR` | `/var/lib/vm-agent/resource-history` | Node-local retry spool for resource-history chunks | +| `RESOURCE_HISTORY_SPOOL_MAX_BYTES` | `20971520` | Max node-local resource-history retry spool bytes | +| `RESOURCE_HISTORY_UPLOAD_TIMEOUT` | `10s` | Deadline for one resource-history upload callback | +| `RESOURCE_HISTORY_MAX_SAMPLES` | `4096` | Max resource samples packed into one uploaded chunk | +| `WORKSPACE_CALLBACK_TOKEN_REFRESH_RATIO` | `0.5` | Fraction of a workspace callback token's lifetime after which the agent renews it (clamped to 0.1–0.9) | +| `WORKSPACE_CALLBACK_TOKEN_RENEWAL_TIMEOUT` | `15s` | Timeout for one workspace callback token renewal request | +| `WORKSPACE_CALLBACK_TOKEN_RENEWAL_RETRY_INITIAL` | `1m` | First backoff after a transient renewal failure | +| `WORKSPACE_CALLBACK_TOKEN_RENEWAL_RETRY_MAX` | `30m` | Renewal backoff ceiling; also the wait after the control plane answers "not yet due" | +| `MSG_AUTH_RENEWAL_WAIT` | `15m` | How long chat-message delivery may stay paused on a rejected workspace token before the pause is reported as an error; held messages are kept | ### Log Retrieval Settings diff --git a/tasks/active/2026-10-04-workspace-callback-token-renewal.md b/tasks/active/2026-10-04-workspace-callback-token-renewal.md index 182079f324..58e0f3d8a6 100644 --- a/tasks/active/2026-10-04-workspace-callback-token-renewal.md +++ b/tasks/active/2026-10-04-workspace-callback-token-renewal.md @@ -103,40 +103,49 @@ gets `401 Invalid or expired callback token` on every workspace-scoped callback. ## Implementation checklist ### API -- [ ] `jwt.ts`: renewal signing preserves generation (`gen` claim); payload exposes generation; - stale-callback guard reads generation before `iat` -- [ ] New service `workspace-callback-token-renewal.ts`: dual-proof verification, D1 binding checks, - due check, Instant superseded check, mint; hibernate-delivery mint helper (active + bound) -- [ ] New callback route file `routes/workspaces/callback-token-renewal.ts` mounted in - `routes/workspaces/index.ts`; `Cache-Control: no-store`; designed 401/403/410; IDs-only logs -- [ ] `node-agent-session-snapshots.ts`: include fresh token on hibernate when bound + active -- [ ] Env vars: `WORKSPACE_CALLBACK_TOKEN_RENEWAL_*` documented (`env.ts`, `.env.example`, docs) +- [x] `jwt.ts`: renewal signing preserves generation (`gen_iat`); claim readers live in + `callback-token-claims.ts` (tests partially mock `jwt.ts`); stale guard reads `gen_iat` first +- [x] New service `workspace-callback-token-renewal.ts`: dual-proof verification, D1 binding checks + (node, owner, active, non-terminal node), due check, JIT re-check, mint; VM-only delivery mint + with post-sign incarnation re-read. Superseded-generation refusal dropped: `agent_sessions.updated_at` + has non-recovery writers (credential attribution, suspend/resume), so it would refuse the live + container; `gen_iat` preservation keeps the stale guard intact instead. +- [x] New callback route file `routes/workspaces/callback-token-renewal.ts` mounted in + `routes/workspaces/index.ts`; node proof in the JSON body (Workers Logs records headers; custom + header redaction is undocumented); `Cache-Control: no-store`; designed 401/403/410; IDs-only logs +- [x] `node-agent-session-snapshots.ts`: include fresh token on hibernate when bound + active (VM only) +- [x] Env vars documented: env-reference skill (API `CALLBACK_TOKEN_*`, agent `WORKSPACE_CALLBACK_TOKEN_*`, + `MSG_AUTH_RENEWAL_WAIT`), public VM agent reference ### VM agent -- [ ] Config: renewal ratio, retry initial/max, request timeout; heartbeat parse unaffected -- [ ] `workspace_callback_token_renewal.go`: due selection with injected clock, request with both - tokens, response classification, bounded backoff, rejection latch, CAS apply -- [ ] Hook renewal after successful heartbeat (TryLock, like ready/eviction retries) -- [ ] `upsertWorkspaceRuntime`: persist token changes and propagate to reporter + session hosts -- [ ] `acp.SessionHost`: lock-free current-token accessor + `SetCallbackToken`; replace reads -- [ ] `messagereport`: stale-token retry, park-on-401, resume on new token, bounded park budget +- [x] Config: renewal ratio (clamped), retry initial/max, request timeout +- [x] `workspace_callback_token_renewal.go`: due selection with injected clock, request with both + tokens, response classification, bounded backoff, rejection latch keyed on the token, CAS apply +- [x] Hook renewal after successful heartbeat (TryLock, like ready/eviction retries) +- [x] `upsertWorkspaceRuntime`: never adopt an earlier-expiring token; persist then propagate +- [x] `acp.SessionHost`: lock-free current-token accessor + `SetCallbackToken`; replace reads +- [x] `messagereport`: stale-token retry, park-on-401 without deleting rows, resume on new token, + pause beyond `MSG_AUTH_RENEWAL_WAIT` reported once (slog.Error + node errorreport) ### Tests -- [ ] Workers test through the real route: renew success (gen preserved, scope workspace, new exp), - expired token denied, node token as workspace proof denied, workspace token as node proof - denied, foreign node denied, workspace on other node (moved) denied, user mismatch denied, - deleted/stopped/evicted workspace 410, terminal node 410, claim/path mismatch, not-due no mint, - Instant superseded denied, concurrent renewals -- [ ] Unit: hibernate push includes token only when bound + active; stale guard uses `gen` -- [ ] Go: clock crosses threshold/24h; expired not sent; rejection latch; backoff; CAS vs concurrent - push; heartbeat (node) vs workspace token separation; propagation to reporter + session host; - persistence across restart; reporter park/resume/no-duplicate/no-loss/budget; real HTTP path -- [ ] `go test -race` for the touched packages +- [x] Workers test through the real route (29): renew success (gen preserved, scope, new exp), + expired denied, node-as-workspace and workspace-as-node proofs denied, foreign/moved node, + owner mismatch, deleted/stopped/terminal node 410, missing row 410, claim/path mismatch, not-due, + concurrent renewals, downstream route accepts the renewed token, VM-only delivery, hibernate body +- [x] Unit (14, real SQLite): delete/move/rebind between reads, Instant excluded, D1 throw → null, + deletion during renewal mint → 410, legacy unscoped proofs, exp boundary, ratio clamping; + stale guard uses `gen_iat` with the real signer +- [x] Go: clock crosses threshold/24h; expired offered once; latch per token; backoff; node-credential + retry; CAS vs concurrent delivery; earlier-expiry delivery ignored; node token untouched; + propagation to reporter + every host of the workspace only; restart hydration; reporter + park/resume/stale-401/no-duplicate/budget/bounded; hibernate handler uses delivered token +- [x] `go test -race` for server, messagereport, acp, config ### Docs / rollout -- [ ] Public docs: security architecture (callback token lifetime + renewal), env reference -- [ ] Rollout note: API push heals old agents for snapshot calls only; other consumers need the new - agent (new nodes); no hot replacement; AI proxy env token residual risk → SAM Idea +- [x] Public docs: security architecture (callback token scopes, lifetime, renewal), VM agent env +- [x] Rollout note (PR body): API push heals old agents for snapshot calls only (proven against pinned + 7a9782c90 source); other consumers need new nodes; no hot replacement; AI proxy env token + residual → SAM Idea 01M432G3276YZWCP3HEJ5B25J5 ## Acceptance criteria @@ -157,6 +166,15 @@ workers tests through the real route plus Go tests against an HTTP control plane deterministic clock-injected Go tests, Miniflare workers tests through the real auth boundary, `-race` runs, and post-deploy production log checks. +## Discrimination evidence + +Each guard was removed once and the intended tests went red, then restored: +API M1 node binding, M2 owner binding, M3 node proof, M4 not-due gate, M5 gen preservation, +M6 claim/path, M7 delivery status, M8 post-sign re-read, M9 Instant exclusion, M10 renewal JIT. +Agent A1 expiry ordering, A2 CAS, A3 refusal latch, A4/A5 propagation, A6 persistence, A7 node +credential retry; reporter R1/R2/R4/R5/R6; SessionHost S1; upsert publish U1. R3 (resume bookkeeping) +is not a guard: parking is keyed on the rejected token, so a new token resumes by construction. + ## References - `.claude/rules/28-credential-resolution-fallback-tests.md`, `.claude/rules/34` (callback auth), From 3e870b5f33f5f86a942ba343420cbc24714c53fb Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 09:06:16 +0000 Subject: [PATCH 07/19] task: record final renewal design decisions Co-Authored-By: Claude Opus 5.5 --- ...-10-04-workspace-callback-token-renewal.md | 19 +++++++++++-------- 1 file changed, 11 insertions(+), 8 deletions(-) diff --git a/tasks/active/2026-10-04-workspace-callback-token-renewal.md b/tasks/active/2026-10-04-workspace-callback-token-renewal.md index 58e0f3d8a6..911e4938f0 100644 --- a/tasks/active/2026-10-04-workspace-callback-token-renewal.md +++ b/tasks/active/2026-10-04-workspace-callback-token-renewal.md @@ -77,15 +77,18 @@ gets `401 Invalid or expired callback token` on every workspace-scoped callback. 1. **API push on hibernate (heals running old agents, no binary rollout needed):** `hibernateAgentSessionOnNode` mints a fresh workspace token only when D1 confirms the workspace - is active on the target node, and sends it as `workspaceCallbackToken` over the existing + is active on the target VM node (never Instant; re-read after signing), and sends it as + `workspaceCallbackToken` over the existing node-management channel (same channel that delivers the token at create/restore). Agents since 2026-07-11 already store it before capture. 2. **Proactive dual-proof renewal (new agents):** `POST /api/workspaces/:id/callback-token/renew` requires the current, unexpired workspace token (`Authorization`) **and** the hosting node's - node token (`X-SAM-Node-ID` / `X-SAM-Node-Authorization`). It renews only if the workspace is - active, bound to that node, owned by the node's user, the node is non-terminal, and (Instant) the - token generation is not superseded. A renewed token preserves the generation issue time (`gen` - claim) so the stale-callback guard keeps working. Not-yet-due tokens are not re-minted. + node token (JSON body `{ nodeId, nodeToken }`; Workers Logs records request headers, and + redaction of custom headers is undocumented). It renews only if the workspace is active, bound to + that node, owned by the node's user, and the node is non-terminal. A renewed token preserves the + generation issue time (`gen_iat` claim) so the Instant stale-callback guard keeps working. No + superseded-generation refusal: its only signal (`agent_sessions.updated_at`) has non-recovery + writers and would refuse the live container. Not-yet-due tokens are not re-minted. - A node token alone cannot obtain a workspace token (no widened node authority). - A workspace token leaked from a VM devcontainer cannot renew itself (needs the node token). - Expired tokens are never renewed (no expiry bypass); recovery is the control-plane push. @@ -94,9 +97,9 @@ gets `401 Invalid or expired callback token` on every workspace-scoped callback. definitive rejections, compare-and-swap the runtime token, persist it, and propagate every token change (renewal or control-plane push) to the message reporter and ACP session hosts. 4. **Message reporter:** a 401 is no longer terminal-and-destructive. If a newer token exists, retry - with it; otherwise keep the outbox, send nothing, and resume when a new token arrives. Bounded by - the existing outbox cap and a new configurable park budget, after which the old terminal - behaviour applies (rule 54.13). + with it; otherwise keep the outbox, send nothing for that token, and resume when a new token + arrives. Bounded by the existing outbox cap; a pause longer than `MSG_AUTH_RENEWAL_WAIT` is + reported once (rule 54.13: no request loop on the rejected token, no row destruction). 5. **ACP SessionHost:** read the callback token through a lock-free accessor that renewal updates (rule 46: nothing reachable from the ACP notification goroutine may take `mu`). From 7b33d4bb2fe5f5b7f16855c6e751278f04b01a24 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 09:18:42 +0000 Subject: [PATCH 08/19] refactor(api): keep the hibernate token delivery free of route helpers The delivery mint pulled routes/workspaces/_helpers into node-agent's import graph, and through it auth.ts, which broke seven unit suites that partially mock their dependencies. The workspace callback identity helpers move to services/workspace-callback-identity.ts (re-exported unchanged from the route helpers), and the binding loader plus delivery mint move to services/workspace-callback-token-binding.ts. Only the renewal route path still uses the route helpers. Co-Authored-By: Claude Opus 5.5 --- apps/api/src/routes/workspaces/_helpers.ts | 50 +++---- .../services/node-agent-session-snapshots.ts | 2 +- .../services/workspace-callback-identity.ts | 38 +++++ .../workspace-callback-token-binding.ts | 139 ++++++++++++++++++ .../workspace-callback-token-renewal.ts | 129 +--------------- .../workspace-callback-token-renewal.test.ts | 6 +- .../workspace-callback-token-renewal.test.ts | 2 +- 7 files changed, 207 insertions(+), 159 deletions(-) create mode 100644 apps/api/src/services/workspace-callback-identity.ts create mode 100644 apps/api/src/services/workspace-callback-token-binding.ts diff --git a/apps/api/src/routes/workspaces/_helpers.ts b/apps/api/src/routes/workspaces/_helpers.ts index d99e7e7344..a9e0463d54 100644 --- a/apps/api/src/routes/workspaces/_helpers.ts +++ b/apps/api/src/routes/workspaces/_helpers.ts @@ -15,32 +15,28 @@ import { signalWorkspaceDeletionUnconfirmedCallback, type WorkspaceDeletionCallbackKind, } from '../../services/workspace-deletion-callback-signal'; +import { + sameWorkspaceCallbackIdentity, + WORKSPACE_CALLBACK_ACTIVE_STATUSES, + type WorkspaceCallbackIdentitySnapshot, +} from '../../services/workspace-callback-identity'; import { resolveWorkspaceGitSource, type WorkspaceGitSourceProject, } from '../../services/workspace-git-source'; +export { + sameWorkspaceCallbackIdentity, + WORKSPACE_CALLBACK_ACTIVE_STATUSES, + type WorkspaceCallbackIdentitySnapshot, +} from '../../services/workspace-callback-identity'; + export const ACTIVE_WORKSPACE_STATUSES = new Set(['running', 'recovery'] as const); -export const WORKSPACE_CALLBACK_ACTIVE_STATUSES: ReadonlySet = new Set([ - 'creating', - 'running', - 'recovery', -]); export const WORKSPACE_CALLBACK_PROVISIONING_FAILURE_STATUSES: ReadonlySet = new Set([ 'creating', 'error', ]); -export interface WorkspaceCallbackIdentitySnapshot { - workspaceId: string; - userId: string; - projectId: string | null; - chatSessionId: string | null; - status: string; - nodeId: string | null; - nodeStatus: string | null; -} - export function isActiveWorkspaceStatus(status: string): boolean { return ACTIVE_WORKSPACE_STATUSES.has(status as 'running' | 'recovery'); } @@ -117,21 +113,6 @@ export async function loadWorkspaceCallbackIdentity( return rows[0] ?? null; } -export function sameWorkspaceCallbackIdentity( - current: WorkspaceCallbackIdentitySnapshot, - expected: WorkspaceCallbackIdentitySnapshot -): boolean { - return ( - current.workspaceId === expected.workspaceId && - current.userId === expected.userId && - current.projectId === expected.projectId && - current.chatSessionId === expected.chatSessionId && - current.status === expected.status && - current.nodeId === expected.nodeId && - current.nodeStatus === expected.nodeStatus - ); -} - interface WorkspaceCallbackTransitionValues { status: string; updatedAt: string; @@ -448,8 +429,13 @@ export async function scheduleWorkspaceCreateOnNode( }, { beforeExternalMutation: assertCurrent } ); - if (options.durableRetry && (!acknowledgement || typeof acknowledgement !== 'object' - || !('workspaceId' in acknowledgement) || acknowledgement.workspaceId !== workspaceId)) { + if ( + options.durableRetry && + (!acknowledgement || + typeof acknowledgement !== 'object' || + !('workspaceId' in acknowledgement) || + acknowledgement.workspaceId !== workspaceId) + ) { throw new Error('Node agent did not acknowledge the expected workspace identity'); } await assertCurrent(); diff --git a/apps/api/src/services/node-agent-session-snapshots.ts b/apps/api/src/services/node-agent-session-snapshots.ts index c75e03edcc..25784a3410 100644 --- a/apps/api/src/services/node-agent-session-snapshots.ts +++ b/apps/api/src/services/node-agent-session-snapshots.ts @@ -5,7 +5,7 @@ import { type GuardedNodeAgentMutationOptions, nodeAgentRequest, } from './node-agent'; -import { mintWorkspaceCallbackTokenForNodeDelivery } from './workspace-callback-token-renewal'; +import { mintWorkspaceCallbackTokenForNodeDelivery } from './workspace-callback-token-binding'; export const DEFAULT_SESSION_SNAPSHOT_REQUEST_TIMEOUT_MS = 5 * 60 * 1000; diff --git a/apps/api/src/services/workspace-callback-identity.ts b/apps/api/src/services/workspace-callback-identity.ts new file mode 100644 index 0000000000..41e1362676 --- /dev/null +++ b/apps/api/src/services/workspace-callback-identity.ts @@ -0,0 +1,38 @@ +/** + * Workspace callback identity: the D1 facts a VM-agent workspace callback is + * authorized against. Dependency-free so lightweight callers (for example the + * hibernate token delivery in node-agent-session-snapshots.ts) can share them + * without importing the route helpers that re-export them + * (routes/workspaces/_helpers.ts). + */ + +export const WORKSPACE_CALLBACK_ACTIVE_STATUSES: ReadonlySet = new Set([ + 'creating', + 'running', + 'recovery', +]); + +export interface WorkspaceCallbackIdentitySnapshot { + workspaceId: string; + userId: string; + projectId: string | null; + chatSessionId: string | null; + status: string; + nodeId: string | null; + nodeStatus: string | null; +} + +export function sameWorkspaceCallbackIdentity( + current: WorkspaceCallbackIdentitySnapshot, + expected: WorkspaceCallbackIdentitySnapshot +): boolean { + return ( + current.workspaceId === expected.workspaceId && + current.userId === expected.userId && + current.projectId === expected.projectId && + current.chatSessionId === expected.chatSessionId && + current.status === expected.status && + current.nodeId === expected.nodeId && + current.nodeStatus === expected.nodeStatus + ); +} diff --git a/apps/api/src/services/workspace-callback-token-binding.ts b/apps/api/src/services/workspace-callback-token-binding.ts new file mode 100644 index 0000000000..c23c8a3196 --- /dev/null +++ b/apps/api/src/services/workspace-callback-token-binding.ts @@ -0,0 +1,139 @@ +/** + * Workspace callback token binding and control-plane delivery. + * + * A workspace callback token may only ever reach the VM node that D1 binds the + * workspace to. This module reads that binding and mints the fresh token the + * control plane delivers over the node-management channel (VM hibernate requests, + * `node-agent-session-snapshots.ts`), exactly as workspace creation does. Instant + * (cf-container) runtimes are excluded: their container DO mints a fresh token on + * every cold wake, and a wall-clock token pushed to a container generation would + * defeat the Instant stale-callback guard (`routes/_stale-callback-guard.ts`). + * + * Kept free of route helpers so that the hibernate path stays lightweight; the + * agent-initiated renewal route lives in `workspace-callback-token-renewal.ts`. + */ +import type { Env } from '../env'; +import { log } from '../lib/logger'; +import { signCallbackToken } from './jwt'; +import { nodeStatusTerminatesCallbacks } from './node-callback-auth'; +import { + sameWorkspaceCallbackIdentity, + WORKSPACE_CALLBACK_ACTIVE_STATUSES, + type WorkspaceCallbackIdentitySnapshot, +} from './workspace-callback-identity'; + +export interface WorkspaceCallbackTokenBinding extends WorkspaceCallbackIdentitySnapshot { + nodeUserId: string | null; + nodeRuntime: string | null; +} + +export async function loadWorkspaceCallbackTokenBinding( + env: Env, + workspaceId: string +): Promise { + const row = await env.DATABASE.prepare( + `SELECT w.id AS workspaceId, + w.user_id AS userId, + w.project_id AS projectId, + w.chat_session_id AS chatSessionId, + w.status AS status, + w.node_id AS nodeId, + n.status AS nodeStatus, + n.user_id AS nodeUserId, + n.runtime AS nodeRuntime + FROM workspaces w + LEFT JOIN nodes n ON n.id = w.node_id + WHERE w.id = ? + LIMIT 1` + ) + .bind(workspaceId) + .first(); + return row ?? null; +} + +/** + * A workspace's callback authority belongs to the node D1 binds it to. Placement only + * binds a workspace to a node owned by the workspace user, so a mismatch is a foreign + * node and must not receive the workspace's credential. + */ +export function workspaceBoundToNode( + binding: WorkspaceCallbackTokenBinding, + nodeId: string +): boolean { + return ( + !!binding.nodeId && + binding.nodeId === nodeId && + !!binding.nodeUserId && + binding.nodeUserId === binding.userId + ); +} + +/** + * Why a binding may not receive a delivered token, or null when it may. Delivery is + * VM-only (see the file header for why Instant runtimes are excluded). + */ +function deliverySkipReason( + binding: WorkspaceCallbackTokenBinding | null, + nodeId: string +): string | null { + if (!binding) return 'workspace_missing'; + if (!workspaceBoundToNode(binding, nodeId)) return 'not_bound_to_node'; + if (binding.nodeRuntime === 'cf-container') return 'instant_runtime'; + if (!WORKSPACE_CALLBACK_ACTIVE_STATUSES.has(binding.status)) return 'workspace_inactive'; + if (!binding.nodeStatus || nodeStatusTerminatesCallbacks(binding.nodeStatus)) { + return 'node_inactive'; + } + return null; +} + +function sameRenewalBinding( + current: WorkspaceCallbackTokenBinding, + expected: WorkspaceCallbackTokenBinding +) { + return ( + sameWorkspaceCallbackIdentity(current, expected) && + current.nodeUserId === expected.nodeUserId && + current.nodeRuntime === expected.nodeRuntime + ); +} + +/** + * Mint a fresh workspace callback token for a control-plane request that is about to be + * delivered to `nodeId` over the node-management channel. Returns null, and the caller + * sends its request without a token exactly as before, unless D1 binds the workspace to + * that VM node and the workspace and node are still active, both before and after + * signing (rule 49), so a delete or move that wins the race gets no credential. + */ +export async function mintWorkspaceCallbackTokenForNodeDelivery( + env: Env, + input: { workspaceId: string; nodeId: string } +): Promise { + const skip = (reason: string) => { + log.info('workspace_callback_token.delivery_skipped', { + workspaceId: input.workspaceId, + nodeId: input.nodeId, + reason, + }); + return null; + }; + try { + const binding = await loadWorkspaceCallbackTokenBinding(env, input.workspaceId); + const skipReason = deliverySkipReason(binding, input.nodeId); + if (skipReason || !binding) return skip(skipReason ?? 'workspace_missing'); + + const token = await signCallbackToken(input.workspaceId, env); + + const current = await loadWorkspaceCallbackTokenBinding(env, input.workspaceId); + if (!current || !sameRenewalBinding(current, binding)) { + return skip('incarnation_changed'); + } + return token; + } catch (err) { + log.warn('workspace_callback_token.delivery_mint_failed', { + workspaceId: input.workspaceId, + nodeId: input.nodeId, + error: err instanceof Error ? err.message : String(err), + }); + return null; + } +} diff --git a/apps/api/src/services/workspace-callback-token-renewal.ts b/apps/api/src/services/workspace-callback-token-renewal.ts index 915fd279de..fec717539d 100644 --- a/apps/api/src/services/workspace-callback-token-renewal.ts +++ b/apps/api/src/services/workspace-callback-token-renewal.ts @@ -13,12 +13,9 @@ * relay (`verifySessionSnapshotRelayAuthorization`). A node token cannot mint a * workspace token, a workspace token copied out of a devcontainer cannot renew * itself, and an expired token is never renewed. - * 2. Control-plane delivery (API -> VM agent, `mintWorkspaceCallbackTokenForNodeDelivery`): - * requests the control plane already sends over the node-management channel to a VM - * node carry a freshly minted token, exactly as workspace create does. Instant - * (cf-container) runtimes are excluded: their container DO mints a fresh token on every - * cold wake, and a wall-clock token pushed to a container generation would defeat the - * Instant stale-callback guard (`routes/_stale-callback-guard.ts`). + * 2. Control-plane delivery (API -> VM agent): requests the control plane already sends + * over the node-management channel to a VM node carry a freshly minted token + * (`mintWorkspaceCallbackTokenForNodeDelivery` in `workspace-callback-token-binding.ts`). * * Both paths renew only while D1 says the workspace is active and bound to the node, so * deleting, stopping or reassigning a workspace still ends its callback authority. @@ -29,16 +26,16 @@ import { AppError, errors } from '../middleware/error'; import { assertWorkspaceAcceptsCallback, assertWorkspaceCallbackIdentityCurrent, - sameWorkspaceCallbackIdentity, - WORKSPACE_CALLBACK_ACTIVE_STATUSES, - type WorkspaceCallbackIdentitySnapshot, } from '../routes/workspaces/_helpers'; import { callbackTokenExpiresAtMs, callbackTokenGenerationIssuedAtSeconds, } from './callback-token-claims'; import { shouldRefreshCallbackToken, signCallbackToken, verifyCallbackToken } from './jwt'; -import { nodeStatusTerminatesCallbacks } from './node-callback-auth'; +import { + loadWorkspaceCallbackTokenBinding, + workspaceBoundToNode, +} from './workspace-callback-token-binding'; /** * Error codes for the node credential, distinct from the workspace credential's @@ -48,52 +45,9 @@ import { nodeStatusTerminatesCallbacks } from './node-callback-auth'; export const NODE_CALLBACK_UNAUTHORIZED = 'NODE_CALLBACK_UNAUTHORIZED'; export const NODE_CALLBACK_FORBIDDEN = 'NODE_CALLBACK_FORBIDDEN'; -interface WorkspaceRenewalBinding extends WorkspaceCallbackIdentitySnapshot { - nodeUserId: string | null; - nodeRuntime: string | null; -} - export type WorkspaceCallbackTokenRenewalResult = { renewed: true; token: string; expiresAt: string | null } | { renewed: false }; -async function loadWorkspaceRenewalBinding( - env: Env, - workspaceId: string -): Promise { - const row = await env.DATABASE.prepare( - `SELECT w.id AS workspaceId, - w.user_id AS userId, - w.project_id AS projectId, - w.chat_session_id AS chatSessionId, - w.status AS status, - w.node_id AS nodeId, - n.status AS nodeStatus, - n.user_id AS nodeUserId, - n.runtime AS nodeRuntime - FROM workspaces w - LEFT JOIN nodes n ON n.id = w.node_id - WHERE w.id = ? - LIMIT 1` - ) - .bind(workspaceId) - .first(); - return row ?? null; -} - -/** - * A workspace's callback authority belongs to the node D1 binds it to. Placement only - * binds a workspace to a node owned by the workspace user, so a mismatch is a foreign - * node and must not receive the workspace's credential. - */ -function workspaceBoundToNode(binding: WorkspaceRenewalBinding, nodeId: string): boolean { - return ( - !!binding.nodeId && - binding.nodeId === nodeId && - !!binding.nodeUserId && - binding.nodeUserId === binding.userId - ); -} - async function verifyRenewalNodeCredential( env: Env, nodeId: string, @@ -144,7 +98,7 @@ export async function renewWorkspaceCallbackToken( } await verifyRenewalNodeCredential(env, nodeId, input.nodeToken); - const binding = await loadWorkspaceRenewalBinding(env, workspaceId); + const binding = await loadWorkspaceCallbackTokenBinding(env, workspaceId); if (binding && !workspaceBoundToNode(binding, nodeId)) { log.warn('workspace_callback_token.renewal_rejected', { workspaceId, @@ -190,70 +144,3 @@ export async function renewWorkspaceCallbackToken( expiresAt: expiresAtMs === null ? null : new Date(expiresAtMs).toISOString(), }; } - -/** - * Why a binding may not receive a delivered token, or null when it may. Delivery is - * VM-only (see the file header for why Instant runtimes are excluded). - */ -function deliverySkipReason( - binding: WorkspaceRenewalBinding | null, - nodeId: string -): string | null { - if (!binding) return 'workspace_missing'; - if (!workspaceBoundToNode(binding, nodeId)) return 'not_bound_to_node'; - if (binding.nodeRuntime === 'cf-container') return 'instant_runtime'; - if (!WORKSPACE_CALLBACK_ACTIVE_STATUSES.has(binding.status)) return 'workspace_inactive'; - if (!binding.nodeStatus || nodeStatusTerminatesCallbacks(binding.nodeStatus)) { - return 'node_inactive'; - } - return null; -} - -function sameRenewalBinding(current: WorkspaceRenewalBinding, expected: WorkspaceRenewalBinding) { - return ( - sameWorkspaceCallbackIdentity(current, expected) && - current.nodeUserId === expected.nodeUserId && - current.nodeRuntime === expected.nodeRuntime - ); -} - -/** - * Mint a fresh workspace callback token for a control-plane request that is about to be - * delivered to `nodeId` over the node-management channel. Returns null, and the caller - * sends its request without a token exactly as before, unless D1 binds the workspace to - * that VM node and the workspace and node are still active, both before and after - * signing (rule 49), so a delete or move that wins the race gets no credential. - */ -export async function mintWorkspaceCallbackTokenForNodeDelivery( - env: Env, - input: { workspaceId: string; nodeId: string } -): Promise { - const skip = (reason: string) => { - log.info('workspace_callback_token.delivery_skipped', { - workspaceId: input.workspaceId, - nodeId: input.nodeId, - reason, - }); - return null; - }; - try { - const binding = await loadWorkspaceRenewalBinding(env, input.workspaceId); - const skipReason = deliverySkipReason(binding, input.nodeId); - if (skipReason || !binding) return skip(skipReason ?? 'workspace_missing'); - - const token = await signCallbackToken(input.workspaceId, env); - - const current = await loadWorkspaceRenewalBinding(env, input.workspaceId); - if (!current || !sameRenewalBinding(current, binding)) { - return skip('incarnation_changed'); - } - return token; - } catch (err) { - log.warn('workspace_callback_token.delivery_mint_failed', { - workspaceId: input.workspaceId, - nodeId: input.nodeId, - error: err instanceof Error ? err.message : String(err), - }); - return null; - } -} diff --git a/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts b/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts index 9f93b7ded4..a6d3cab94e 100644 --- a/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts +++ b/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts @@ -16,10 +16,8 @@ import type { Env } from '../../../src/env'; import { AppError } from '../../../src/middleware/error'; import { CALLBACK_TOKEN_GENERATION_ISSUED_AT_CLAIM } from '../../../src/services/callback-token-claims'; import { signNodeCallbackToken } from '../../../src/services/jwt'; -import { - mintWorkspaceCallbackTokenForNodeDelivery, - renewWorkspaceCallbackToken, -} from '../../../src/services/workspace-callback-token-renewal'; +import { mintWorkspaceCallbackTokenForNodeDelivery } from '../../../src/services/workspace-callback-token-binding'; +import { renewWorkspaceCallbackToken } from '../../../src/services/workspace-callback-token-renewal'; import { createSchemaTables, createSqliteD1 } from '../../helpers/sqlite-d1'; const HOUR = 3600; diff --git a/apps/api/tests/workers/workspace-callback-token-renewal.test.ts b/apps/api/tests/workers/workspace-callback-token-renewal.test.ts index a320cd30df..52aaa99832 100644 --- a/apps/api/tests/workers/workspace-callback-token-renewal.test.ts +++ b/apps/api/tests/workers/workspace-callback-token-renewal.test.ts @@ -23,7 +23,7 @@ import { verifyCallbackToken, } from '../../src/services/jwt'; import { hibernateAgentSessionOnNode } from '../../src/services/node-agent-session-snapshots'; -import { mintWorkspaceCallbackTokenForNodeDelivery } from '../../src/services/workspace-callback-token-renewal'; +import { mintWorkspaceCallbackTokenForNodeDelivery } from '../../src/services/workspace-callback-token-binding'; import { seedNode, seedUser, seedWorkspace } from './helpers/seed-d1'; const testEnv = env as unknown as Env; From b80a7796bef29753c6b04d126631265931aea620 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 09:18:59 +0000 Subject: [PATCH 09/19] style(api): sort route helper exports Co-Authored-By: Claude Opus 5.5 --- apps/api/src/routes/workspaces/_helpers.ts | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/apps/api/src/routes/workspaces/_helpers.ts b/apps/api/src/routes/workspaces/_helpers.ts index a9e0463d54..09e2515fbf 100644 --- a/apps/api/src/routes/workspaces/_helpers.ts +++ b/apps/api/src/routes/workspaces/_helpers.ts @@ -11,15 +11,15 @@ import { errors } from '../../middleware/error'; import { signCallbackToken, verifyCallbackToken } from '../../services/jwt'; import { createWorkspaceOnNode } from '../../services/node-agent'; import { nodeStatusTerminatesCallbacks } from '../../services/node-callback-auth'; -import { - signalWorkspaceDeletionUnconfirmedCallback, - type WorkspaceDeletionCallbackKind, -} from '../../services/workspace-deletion-callback-signal'; import { sameWorkspaceCallbackIdentity, WORKSPACE_CALLBACK_ACTIVE_STATUSES, type WorkspaceCallbackIdentitySnapshot, } from '../../services/workspace-callback-identity'; +import { + signalWorkspaceDeletionUnconfirmedCallback, + type WorkspaceDeletionCallbackKind, +} from '../../services/workspace-deletion-callback-signal'; import { resolveWorkspaceGitSource, type WorkspaceGitSourceProject, From 8bf5fe7dbf37b52156915d87221960d6178f57c2 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 09:55:38 +0000 Subject: [PATCH 10/19] fix(api): rate-limit workspace token renewal and keep it VM-only Review findings on the renewal route: - Rule 28 requires an atomic per-principal limit on credential rotation endpoints. Authenticated renewal attempts now count against RATE_LIMIT_CALLBACK_TOKEN_RENEWAL (default 12 per hour per workspace) in a new D1 table, with one guarded upsert. A slot is spent only after both proofs and the node binding pass, so a caller without the workspace's credentials cannot use up its quota. Over the limit the route answers 429 with Retry-After; the agent backs off. - Renewal now refuses Instant (cf-container) workspaces, as delivery already did. Their container DO mints a token per cold wake, one per generation, so a superseded generation could otherwise extend its workspace authority. - The sleep wait loop called the hibernate request once per poll and minted a token each time. The agent installs a delivered token whether or not it accepts the request, so one delivered token now serves every poll. Co-Authored-By: Claude Opus 5.5 --- .claude/skills/api-reference/SKILL.md | 2 +- .claude/skills/env-reference/SKILL.md | 2 + apps/api/.env.example | 6 + ...ace_callback_token_renewal_rate_limits.sql | 15 ++ apps/api/src/db/schema.ts | 14 ++ apps/api/src/env.ts | 2 + apps/api/src/middleware/rate-limit.ts | 4 + .../workspaces/callback-token-renewal.ts | 29 +++- .../services/node-agent-session-snapshots.ts | 28 +++- apps/api/src/services/node-agent.ts | 1 + .../services/session-sleep-snapshot-wait.ts | 10 +- .../workspace-callback-token-binding.ts | 23 ++- ...space-callback-token-renewal-rate-limit.ts | 80 +++++++++ .../workspace-callback-token-renewal.ts | 37 ++++- ...-callback-token-renewal-rate-limit.test.ts | 154 ++++++++++++++++++ .../workspace-callback-token-renewal.test.ts | 48 +++++- .../unit/session-sleep-snapshot-wait.test.ts | 69 +++++++- .../workspace-callback-token-renewal.test.ts | 122 +++++++++++++- .../docs/docs/architecture/security.md | 8 +- 19 files changed, 621 insertions(+), 33 deletions(-) create mode 100644 apps/api/src/db/migrations/0179_workspace_callback_token_renewal_rate_limits.sql create mode 100644 apps/api/src/services/workspace-callback-token-renewal-rate-limit.ts create mode 100644 apps/api/tests/unit/services/workspace-callback-token-renewal-rate-limit.test.ts diff --git a/.claude/skills/api-reference/SKILL.md b/.claude/skills/api-reference/SKILL.md index 624c8da8b3..f6875804fe 100644 --- a/.claude/skills/api-reference/SKILL.md +++ b/.claude/skills/api-reference/SKILL.md @@ -220,7 +220,7 @@ The MCP `create_trigger` tool intentionally creates cron triggers only. Generic - `POST /api/workspaces/:id/provisioning-failed` — Workspace provisioning failure callback (sets workspace to `error`) - `POST /api/workspaces/:id/heartbeat` — Workspace activity heartbeat callback - `GET /api/workspaces/:id/runtime` — Workspace runtime metadata callback (repository/branch for recovery) -- `POST /api/workspaces/:id/callback-token/renew` — Renew a workspace callback token before it expires. Two proofs: the current, unexpired workspace token in `Authorization`, and `{ nodeId, nodeToken }` (the hosting node's node-scoped token) in the body. Renews only while the workspace is `creating`/`running`/`recovery`, bound to that node, owned by the node's user, on a non-terminal node, and past `CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO` of the token's lifetime (otherwise `{ renewed: false }`). The renewed token keeps the chain's first `iat` as `gen_iat`. Returns `{ renewed: true, token, expiresAt }` with `Cache-Control: no-store`; 401/403/410 otherwise (`NODE_CALLBACK_UNAUTHORIZED`/`NODE_CALLBACK_FORBIDDEN` when only the node proof failed) +- `POST /api/workspaces/:id/callback-token/renew` — Renew a workspace callback token before it expires. Two proofs: the current, unexpired workspace token in `Authorization`, and `{ nodeId, nodeToken }` (the hosting node's node-scoped token) in the body. The workspace must be `creating`/`running`/`recovery`, bound to that VM node (never an Instant `cf-container` node), owned by the node's user, and on a non-terminal node; otherwise 401/403/410 (`NODE_CALLBACK_UNAUTHORIZED`/`NODE_CALLBACK_FORBIDDEN` when only the node proof failed). Authenticated attempts count against `RATE_LIMIT_CALLBACK_TOKEN_RENEWAL` per workspace per window (429 `RATE_LIMIT_EXCEEDED` with `Retry-After`). A token younger than `CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO` of its lifetime gets `{ renewed: false }`; otherwise `{ renewed: true, token, expiresAt }`, where the token keeps the chain's first `iat` as `gen_iat`. Responses carry `Cache-Control: no-store` - `POST /api/workspaces/:id/boot-log` — Workspace boot progress log callback - `POST /api/workspaces/:id/agent-settings` — Workspace agent settings callback (model, permissionMode) - `POST /api/projects/:id/workspaces/:workspaceId/eviction` — VM-agent callback JWT endpoint that validates node/workspace/runtime-generation identity and successful container stop, atomically marks the workspace `evicted` and closes usage/agent sessions, then serializes replay-safe ProjectData finalization through NodeLifecycle. A stale generation returns 410; failed finalization is retryable. Explicit restart requires renewed capacity admission and rotates the runtime generation diff --git a/.claude/skills/env-reference/SKILL.md b/.claude/skills/env-reference/SKILL.md index 6d44dcfb72..e0b1837525 100644 --- a/.claude/skills/env-reference/SKILL.md +++ b/.claude/skills/env-reference/SKILL.md @@ -538,6 +538,8 @@ by the read-only cron-liveness check. - `CALLBACK_TOKEN_EXPIRY_MS` — Lifetime of node- and workspace-scoped VM callback JWTs (default: `86400000` / 24h). Changing it does not extend tokens already issued. - `CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO` — Fraction of a callback token's lifetime after which it may be renewed (default: `0.5`, clamped to `0.1`–`0.9`). Gates both the node token refresh in `POST /api/nodes/:id/heartbeat` and workspace token renewal in `POST /api/workspaces/:id/callback-token/renew`; a token younger than this is not re-minted. +- `RATE_LIMIT_CALLBACK_TOKEN_RENEWAL` — Authenticated workspace callback-token renewal attempts allowed per workspace per window (default: `12`). Counted atomically in D1 (`workspace_callback_token_renewal_rate_limits`) only after both proofs and the node binding pass; a healthy agent asks about once per half token lifetime. +- `RATE_LIMIT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS` — Window for `RATE_LIMIT_CALLBACK_TOKEN_RENEWAL` (default: `3600`). ### Timeouts diff --git a/apps/api/.env.example b/apps/api/.env.example index e93772fb76..47c16e3d9a 100644 --- a/apps/api/.env.example +++ b/apps/api/.env.example @@ -492,6 +492,12 @@ DEBUG_AGENT_MODEL_OUTPUT_TOKENS=4096 # MAX_CLIENT_ERROR_BATCH_SIZE=25 # MAX_CLIENT_ERROR_BODY_BYTES=65536 +# VM agent callback tokens (node- and workspace-scoped callback JWTs) +# CALLBACK_TOKEN_EXPIRY_MS=86400000 # Token lifetime (default: 24h) +# CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO=0.5 # Lifetime fraction before renewal (default: 0.5, clamped 0.1-0.9) +# RATE_LIMIT_CALLBACK_TOKEN_RENEWAL=12 # Workspace token renewal attempts per workspace per window (default: 12) +# RATE_LIMIT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS=3600 # Renewal rate-limit window in seconds (default: 3600) + # VM agent error reporting # MAX_VM_AGENT_ERROR_BODY_BYTES=32768 # MAX_VM_AGENT_ERROR_BATCH_SIZE=10 diff --git a/apps/api/src/db/migrations/0179_workspace_callback_token_renewal_rate_limits.sql b/apps/api/src/db/migrations/0179_workspace_callback_token_renewal_rate_limits.sql new file mode 100644 index 0000000000..2047d7b54c --- /dev/null +++ b/apps/api/src/db/migrations/0179_workspace_callback_token_renewal_rate_limits.sql @@ -0,0 +1,15 @@ +-- Atomic per-workspace counters for workspace callback-token renewal +-- (POST /api/workspaces/:id/callback-token/renew). +-- +-- A credential rotation endpoint needs a per-principal limit whose state cannot +-- lose increments under concurrent requests, which KV read-modify-write cannot +-- guarantee (.claude/rules/28). One row per workspace; a new window replaces the +-- counter in place, and deleting the workspace deletes its row. +-- +-- Additive only: a new table. No DROP, no table rebuild. + +CREATE TABLE workspace_callback_token_renewal_rate_limits ( + workspace_id TEXT PRIMARY KEY NOT NULL REFERENCES workspaces(id) ON DELETE CASCADE, + window_start INTEGER NOT NULL CHECK (window_start >= 0), + count INTEGER NOT NULL CHECK (count >= 0) +); diff --git a/apps/api/src/db/schema.ts b/apps/api/src/db/schema.ts index 9009b3e3d5..a1e9463c49 100644 --- a/apps/api/src/db/schema.ts +++ b/apps/api/src/db/schema.ts @@ -160,6 +160,20 @@ export const aiSpendRateLimits = sqliteTable( }) ); +// ============================================================================= +// Workspace Callback Token Renewal Rate Limits +// ============================================================================= +export const workspaceCallbackTokenRenewalRateLimits = sqliteTable( + 'workspace_callback_token_renewal_rate_limits', + { + workspaceId: text('workspace_id') + .primaryKey() + .references(() => workspaces.id, { onDelete: 'cascade' }), + windowStart: integer('window_start').notNull(), + count: integer('count').notNull(), + } +); + // ============================================================================= // Sessions (BetterAuth) // ============================================================================= diff --git a/apps/api/src/env.ts b/apps/api/src/env.ts index eb34512b78..f3ff6d4d51 100644 --- a/apps/api/src/env.ts +++ b/apps/api/src/env.ts @@ -334,6 +334,8 @@ export interface Env extends WebhookTriggerEnv, TaskRecoveryEnv { RATE_LIMIT_SESSION_SUMMARIZE_WINDOW_SECONDS?: string; RATE_LIMIT_IDENTITY_TOKEN?: string; RATE_LIMIT_IDENTITY_TOKEN_WINDOW_SECONDS?: string; + RATE_LIMIT_CALLBACK_TOKEN_RENEWAL?: string; // Authenticated workspace callback-token renewal attempts per workspace per window (default: 12) + RATE_LIMIT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS?: string; // Window for RATE_LIMIT_CALLBACK_TOKEN_RENEWAL (default: 3600) /** * Max Codex refresh requests per user per window. Defaults to 30. Enforced * atomically by CodexRefreshLock DO using ctx.storage (not KV). See diff --git a/apps/api/src/middleware/rate-limit.ts b/apps/api/src/middleware/rate-limit.ts index f1dddcc701..f0f43a8fae 100644 --- a/apps/api/src/middleware/rate-limit.ts +++ b/apps/api/src/middleware/rate-limit.ts @@ -54,6 +54,10 @@ export const DEFAULT_RATE_LIMITS = { // Voice transcription runs Workers AI Whisper on every request. Per MINUTE, not per hour // (`DEFAULT_TRANSCRIBE_WINDOW_SECONDS`): dictation is bursty, and the budget is for abuse. TRANSCRIBE: 30, + // Workspace callback-token renewal, per workspace. A healthy VM agent asks about once + // per half token lifetime, so this only bounds a holder of both proofs replaying them + // (`services/workspace-callback-token-renewal-rate-limit.ts`). + CALLBACK_TOKEN_RENEWAL: 12, } as const; /** Default time window (1 hour in seconds) */ diff --git a/apps/api/src/routes/workspaces/callback-token-renewal.ts b/apps/api/src/routes/workspaces/callback-token-renewal.ts index 860b81da07..4569952d02 100644 --- a/apps/api/src/routes/workspaces/callback-token-renewal.ts +++ b/apps/api/src/routes/workspaces/callback-token-renewal.ts @@ -6,7 +6,8 @@ * CALLBACK_TOKEN_EXPIRY_MS keeps working. Auth is two callback JWTs, never a session cookie * (.claude/rules/34): the workspace's current token in `Authorization`, and the hosting * node's id and node token in the JSON body. See - * `services/workspace-callback-token-renewal.ts` for the binding rules. + * `services/workspace-callback-token-renewal.ts` for the renewal rules and + * `services/workspace-callback-token-binding.ts` for the node binding they enforce. * * `workspacesRoutes` applies no session middleware, so this callback route is safe to * mount there next to the other workspace callbacks (`/:id/messages`, `/:id/git-token`). @@ -17,7 +18,11 @@ import * as v from 'valibot'; import type { Env } from '../../env'; import { extractBearerToken } from '../../lib/auth-helpers'; import { errors } from '../../middleware/error'; -import { renewWorkspaceCallbackToken } from '../../services/workspace-callback-token-renewal'; +import { RateLimitError } from '../../middleware/rate-limit'; +import { + renewWorkspaceCallbackToken, + type WorkspaceCallbackTokenRenewalResult, +} from '../../services/workspace-callback-token-renewal'; const MAX_NODE_ID_LENGTH = 128; const MAX_NODE_TOKEN_LENGTH = 16 * 1024; @@ -46,12 +51,20 @@ callbackTokenRenewalRoutes.post('/:id/callback-token/renew', async (c) => { throw errors.badRequest('Invalid callback token renewal request'); } - const result = await renewWorkspaceCallbackToken(c.env, { - workspaceId, - workspaceToken, - nodeId: parsed.output.nodeId, - nodeToken: parsed.output.nodeToken, - }); + let result: WorkspaceCallbackTokenRenewalResult; + try { + result = await renewWorkspaceCallbackToken(c.env, { + workspaceId, + workspaceToken, + nodeId: parsed.output.nodeId, + nodeToken: parsed.output.nodeToken, + }); + } catch (err) { + if (err instanceof RateLimitError) { + c.header('Retry-After', Math.max(1, err.retryAfter).toString()); + } + throw err; + } // The response can carry a credential; no intermediary may store it. c.header('Cache-Control', 'no-store'); return c.json(result); diff --git a/apps/api/src/services/node-agent-session-snapshots.ts b/apps/api/src/services/node-agent-session-snapshots.ts index 25784a3410..d32060b7ae 100644 --- a/apps/api/src/services/node-agent-session-snapshots.ts +++ b/apps/api/src/services/node-agent-session-snapshots.ts @@ -57,21 +57,39 @@ function requestSessionSnapshot( ); } +/** Supplies the workspace token one hibernate request delivers; null delivers none. */ +export type HibernateCallbackTokenDelivery = () => Promise; + +/** + * Mint at most one delivered workspace token for repeats of the same hibernate request. + * The agent installs a delivered token as soon as it reads the request, accepted or not, + * so a caller that repeats the request until the agent accepts it resends the token it + * minted first instead of signing a new one on every poll. + */ +export function hibernateCallbackTokenDelivery( + env: Env, + target: { workspaceId: string; nodeId: string } +): HibernateCallbackTokenDelivery { + let minted: Promise | undefined; + return () => (minted ??= mintWorkspaceCallbackTokenForNodeDelivery(env, target)); +} + export async function hibernateAgentSessionOnNode( nodeId: string, workspaceId: string, sessionId: string, env: Env, userId: string, - input: SessionSnapshotRequest + input: SessionSnapshotRequest, + deliverWorkspaceCallbackToken: HibernateCallbackTokenDelivery = hibernateCallbackTokenDelivery( + env, + { workspaceId, nodeId } + ) ): Promise { // The capture's prepare/progress/complete/failure callbacks authenticate with the // workspace token the agent holds. Deliver a fresh one over this node-management // request, exactly as create/restore do, so a long-awake workspace can still sleep. - const workspaceCallbackToken = await mintWorkspaceCallbackTokenForNodeDelivery(env, { - workspaceId, - nodeId, - }); + const workspaceCallbackToken = await deliverWorkspaceCallbackToken(); return requestSessionSnapshot('hibernate', nodeId, workspaceId, sessionId, env, userId, { ...input, ...(workspaceCallbackToken ? { workspaceCallbackToken } : {}), diff --git a/apps/api/src/services/node-agent.ts b/apps/api/src/services/node-agent.ts index 64618b5a60..88a260b796 100644 --- a/apps/api/src/services/node-agent.ts +++ b/apps/api/src/services/node-agent.ts @@ -757,6 +757,7 @@ export async function sendPromptToAgentOnNode( export { hibernateAgentSessionOnNode, + hibernateCallbackTokenDelivery, restoreAgentSessionOnNode, } from './node-agent-session-snapshots'; diff --git a/apps/api/src/services/session-sleep-snapshot-wait.ts b/apps/api/src/services/session-sleep-snapshot-wait.ts index 2eb7a31cea..58ba94135c 100644 --- a/apps/api/src/services/session-sleep-snapshot-wait.ts +++ b/apps/api/src/services/session-sleep-snapshot-wait.ts @@ -4,7 +4,7 @@ import type * as schema from '../db/schema'; import type { Env } from '../env'; import { log } from '../lib/logger'; import { parsePositiveInt } from '../lib/route-helpers'; -import { hibernateAgentSessionOnNode } from './node-agent'; +import { hibernateAgentSessionOnNode, hibernateCallbackTokenDelivery } from './node-agent'; import { completeActiveSessionSnapshotAsDegraded, getSessionSnapshotCaptureState, @@ -67,6 +67,11 @@ export async function waitForFinalSessionSnapshot( let activeCaptureGeneration: string | null = null; let lastProgressAt = Date.now(); let lastProgressToken = ''; + // One delivered workspace token for every poll of this request, not one per poll. + const deliverWorkspaceCallbackToken = hibernateCallbackTokenDelivery(env, { + workspaceId: input.workspaceId, + nodeId: input.nodeId, + }); while (Date.now() < requestDeadline) { const current = await getSessionSnapshotCaptureState(db, input.chatSessionId); @@ -90,7 +95,8 @@ export async function waitForFinalSessionSnapshot( runtime: input.runtime, agentType: input.agentType, background: true, - } + }, + deliverWorkspaceCallbackToken )) as SnapshotResult & { accepted?: unknown }; if (result.status !== 'pending') { throw new Error(`Workspace snapshot request was not accepted (${String(result.status)})`); diff --git a/apps/api/src/services/workspace-callback-token-binding.ts b/apps/api/src/services/workspace-callback-token-binding.ts index c23c8a3196..10bbbf17c1 100644 --- a/apps/api/src/services/workspace-callback-token-binding.ts +++ b/apps/api/src/services/workspace-callback-token-binding.ts @@ -4,10 +4,15 @@ * A workspace callback token may only ever reach the VM node that D1 binds the * workspace to. This module reads that binding and mints the fresh token the * control plane delivers over the node-management channel (VM hibernate requests, - * `node-agent-session-snapshots.ts`), exactly as workspace creation does. Instant - * (cf-container) runtimes are excluded: their container DO mints a fresh token on - * every cold wake, and a wall-clock token pushed to a container generation would - * defeat the Instant stale-callback guard (`routes/_stale-callback-guard.ts`). + * `node-agent-session-snapshots.ts`), exactly as workspace creation does. + * + * Renewal and delivery are VM-only (`isInstantRuntimeBinding`). An Instant + * (cf-container) runtime gets a fresh token from its container DO on every cold wake, + * one per container generation, and recovery replaces a generation under the same + * nodeId. Renewing would let a superseded generation that is still running extend its + * workspace authority past the lifetime of the token it was started with, and a + * wall-clock token pushed to a generation would defeat the Instant stale-callback + * guard (`routes/_stale-callback-guard.ts`). * * Kept free of route helpers so that the hibernate path stays lightweight; the * agent-initiated renewal route lives in `workspace-callback-token-renewal.ts`. @@ -68,9 +73,13 @@ export function workspaceBoundToNode( ); } +/** Instant (cf-container) workspaces neither renew nor receive tokens (see file header). */ +export function isInstantRuntimeBinding(binding: WorkspaceCallbackTokenBinding): boolean { + return binding.nodeRuntime === 'cf-container'; +} + /** - * Why a binding may not receive a delivered token, or null when it may. Delivery is - * VM-only (see the file header for why Instant runtimes are excluded). + * Why a binding may not receive a delivered token, or null when it may. */ function deliverySkipReason( binding: WorkspaceCallbackTokenBinding | null, @@ -78,7 +87,7 @@ function deliverySkipReason( ): string | null { if (!binding) return 'workspace_missing'; if (!workspaceBoundToNode(binding, nodeId)) return 'not_bound_to_node'; - if (binding.nodeRuntime === 'cf-container') return 'instant_runtime'; + if (isInstantRuntimeBinding(binding)) return 'instant_runtime'; if (!WORKSPACE_CALLBACK_ACTIVE_STATUSES.has(binding.status)) return 'workspace_inactive'; if (!binding.nodeStatus || nodeStatusTerminatesCallbacks(binding.nodeStatus)) { return 'node_inactive'; diff --git a/apps/api/src/services/workspace-callback-token-renewal-rate-limit.ts b/apps/api/src/services/workspace-callback-token-renewal-rate-limit.ts new file mode 100644 index 0000000000..cb13d016db --- /dev/null +++ b/apps/api/src/services/workspace-callback-token-renewal-rate-limit.ts @@ -0,0 +1,80 @@ +/** + * Per-workspace limit for POST /api/workspaces/:id/callback-token/renew. + * + * A credential rotation endpoint needs a per-principal limit whose state cannot lose + * increments under concurrent requests (.claude/rules/28), so this is one guarded D1 + * upsert, not the generic KV limiter. The renewal service spends a slot only after both + * proofs verified and D1 confirmed the workspace is active on the calling node, so a + * party without the workspace's credentials cannot use up its quota. + * + * A healthy VM agent asks about once per half token lifetime and waits + * WORKSPACE_CALLBACK_TOKEN_RETRY_MAX after a not-due answer, far below the default. The + * limit bounds a holder of both proofs replaying them in a loop. + */ +import type { Env } from '../env'; +import { parsePositiveInt } from '../lib/route-helpers'; +import { getRateLimit } from '../middleware/rate-limit'; + +/** Window for RATE_LIMIT_CALLBACK_TOKEN_RENEWAL. */ +export const DEFAULT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS = 3600; + +export function getCallbackTokenRenewalRateLimit(env: Env): { + limit: number; + windowSeconds: number; +} { + return { + limit: getRateLimit(env, 'CALLBACK_TOKEN_RENEWAL'), + windowSeconds: parsePositiveInt( + env.RATE_LIMIT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS, + DEFAULT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS + ), + }; +} + +export type CallbackTokenRenewalQuota = + | { outcome: 'allowed' } + | { outcome: 'limited'; retryAfterSeconds: number } + | { outcome: 'workspace_missing' }; + +/** + * Atomically count one renewal attempt for `workspaceId` in the current fixed window. + * + * SQLite serializes the single upsert, so concurrent attempts cannot both read the same + * count. The row is inserted only while the workspace row exists; a workspace deleted + * since the caller's checks returns `workspace_missing` instead of a foreign-key error. + */ +export async function consumeCallbackTokenRenewalQuota( + env: Env, + workspaceId: string, + nowMs = Date.now() +): Promise { + const { limit, windowSeconds } = getCallbackTokenRenewalRateLimit(env); + const nowSeconds = Math.floor(nowMs / 1000); + const windowStart = Math.floor(nowSeconds / windowSeconds) * windowSeconds; + + // The SELECT needs its WHERE clause: SQLite requires one to parse an upsert whose + // values come from a SELECT. + const row = await env.DATABASE.prepare( + `INSERT INTO workspace_callback_token_renewal_rate_limits (workspace_id, window_start, count) + SELECT id, ?, 1 FROM workspaces WHERE id = ? + ON CONFLICT(workspace_id) DO UPDATE SET + count = CASE + WHEN workspace_callback_token_renewal_rate_limits.window_start = excluded.window_start + THEN workspace_callback_token_renewal_rate_limits.count + 1 + ELSE 1 + END, + window_start = excluded.window_start + RETURNING count` + ) + .bind(windowStart, workspaceId) + .first<{ count: number }>(); + + if (!row) return { outcome: 'workspace_missing' }; + if (Number(row.count) > limit) { + return { + outcome: 'limited', + retryAfterSeconds: Math.max(1, windowStart + windowSeconds - nowSeconds), + }; + } + return { outcome: 'allowed' }; +} diff --git a/apps/api/src/services/workspace-callback-token-renewal.ts b/apps/api/src/services/workspace-callback-token-renewal.ts index fec717539d..0311126a62 100644 --- a/apps/api/src/services/workspace-callback-token-renewal.ts +++ b/apps/api/src/services/workspace-callback-token-renewal.ts @@ -18,11 +18,14 @@ * (`mintWorkspaceCallbackTokenForNodeDelivery` in `workspace-callback-token-binding.ts`). * * Both paths renew only while D1 says the workspace is active and bound to the node, so - * deleting, stopping or reassigning a workspace still ends its callback authority. + * deleting, stopping or reassigning a workspace still ends its callback authority. Both + * are VM-only (`isInstantRuntimeBinding`), and renewal attempts are rate limited per + * workspace (`workspace-callback-token-renewal-rate-limit.ts`). */ import type { Env } from '../env'; import { log } from '../lib/logger'; import { AppError, errors } from '../middleware/error'; +import { RateLimitError } from '../middleware/rate-limit'; import { assertWorkspaceAcceptsCallback, assertWorkspaceCallbackIdentityCurrent, @@ -33,9 +36,11 @@ import { } from './callback-token-claims'; import { shouldRefreshCallbackToken, signCallbackToken, verifyCallbackToken } from './jwt'; import { + isInstantRuntimeBinding, loadWorkspaceCallbackTokenBinding, workspaceBoundToNode, } from './workspace-callback-token-binding'; +import { consumeCallbackTokenRenewalQuota } from './workspace-callback-token-renewal-rate-limit'; /** * Error codes for the node credential, distinct from the workspace credential's @@ -72,11 +77,12 @@ async function verifyRenewalNodeCredential( } /** - * Renew a workspace callback token for the agent on the node hosting the workspace. + * Renew a workspace callback token for the agent on the VM node hosting the workspace. * * Returns `{ renewed: false }` while the presented token is younger than the shared * CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO, so a holder cannot mint early. Throws designed - * 401/403/410 AppErrors; never a 5xx for a rejected credential. + * 401/403/410 AppErrors, or a 429 RateLimitError once the workspace has used its renewal + * attempts for the window; never a 5xx for a rejected credential. */ export async function renewWorkspaceCallbackToken( env: Env, @@ -109,6 +115,15 @@ export async function renewWorkspaceCallbackToken( }); throw errors.forbidden('Workspace is not hosted on this node'); } + if (binding && isInstantRuntimeBinding(binding)) { + log.info('workspace_callback_token.renewal_rejected', { + workspaceId, + nodeId, + reason: 'instant_runtime', + action: 'rejected', + }); + throw errors.forbidden('Instant workspaces get a new callback token when they wake'); + } const active = await assertWorkspaceAcceptsCallback( env, binding, @@ -116,6 +131,22 @@ export async function renewWorkspaceCallbackToken( 'callback_token_renewal' ); + // Counted only after both proofs and the binding passed, so a party without the + // workspace's credentials cannot use up its renewal attempts. + const quota = await consumeCallbackTokenRenewalQuota(env, workspaceId); + if (quota.outcome === 'workspace_missing') { + throw errors.gone('Workspace is missing; callback resource is gone'); + } + if (quota.outcome === 'limited') { + log.warn('workspace_callback_token.renewal_rate_limited', { + workspaceId, + nodeId, + retryAfterSeconds: quota.retryAfterSeconds, + action: 'rejected', + }); + throw new RateLimitError(quota.retryAfterSeconds); + } + if (!shouldRefreshCallbackToken(workspaceToken, env)) { return { renewed: false }; } diff --git a/apps/api/tests/unit/services/workspace-callback-token-renewal-rate-limit.test.ts b/apps/api/tests/unit/services/workspace-callback-token-renewal-rate-limit.test.ts new file mode 100644 index 0000000000..eaa3ec34d4 --- /dev/null +++ b/apps/api/tests/unit/services/workspace-callback-token-renewal-rate-limit.test.ts @@ -0,0 +1,154 @@ +/** + * The per-workspace renewal limit against a real SQLite engine (rule 28: a limit whose state + * is a SQL statement must be proven on a SQL engine, not a mock). Route-level behaviour, + * including that only authenticated attempts are counted, is in + * tests/workers/workspace-callback-token-renewal.test.ts. + */ +import Database from 'better-sqlite3'; +import { beforeEach, describe, expect, it } from 'vitest'; + +import * as schema from '../../../src/db/schema'; +import type { Env } from '../../../src/env'; +import { DEFAULT_RATE_LIMITS } from '../../../src/middleware/rate-limit'; +import { + consumeCallbackTokenRenewalQuota, + DEFAULT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS, + getCallbackTokenRenewalRateLimit, +} from '../../../src/services/workspace-callback-token-renewal-rate-limit'; +import { createSchemaTables, createSqliteD1 } from '../../helpers/sqlite-d1'; + +const WS = 'ws-1'; +const OTHER_WS = 'ws-2'; +const WINDOW_SECONDS = 600; +// Mid-window, so a burst never straddles a window boundary by accident. +const T0 = (1_000_000 * WINDOW_SECONDS + 100) * 1000; + +let sqlite: Database.Database; + +function makeEnv(overrides: Record = {}): Env { + return { + DATABASE: createSqliteD1(sqlite), + RATE_LIMIT_CALLBACK_TOKEN_RENEWAL: '3', + RATE_LIMIT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS: String(WINDOW_SECONDS), + ...overrides, + } as unknown as Env; +} + +function counterRows(): Array<{ workspace_id: string; window_start: number; count: number }> { + return sqlite + .prepare('SELECT * FROM workspace_callback_token_renewal_rate_limits ORDER BY workspace_id') + .all() as Array<{ workspace_id: string; window_start: number; count: number }>; +} + +beforeEach(() => { + sqlite = new Database(':memory:'); + createSchemaTables(sqlite, [schema.workspaces, schema.workspaceCallbackTokenRenewalRateLimits]); + const insert = sqlite.prepare( + `INSERT INTO workspaces (id, user_id, node_id, status, name) VALUES (?, 'user-1', 'node-1', 'running', 'ws')` + ); + insert.run(WS); + insert.run(OTHER_WS); +}); + +describe('consumeCallbackTokenRenewalQuota', () => { + it('allows the configured attempts per window, then reports when the window ends', async () => { + const env = makeEnv(); + for (let attempt = 0; attempt < 3; attempt++) { + expect(await consumeCallbackTokenRenewalQuota(env, WS, T0)).toEqual({ outcome: 'allowed' }); + } + + expect(await consumeCallbackTokenRenewalQuota(env, WS, T0 + 1000)).toEqual({ + outcome: 'limited', + retryAfterSeconds: WINDOW_SECONDS - 101, + }); + }); + + it('starts a new count when the window rolls over, and bounds that burst too', async () => { + const env = makeEnv(); + for (let attempt = 0; attempt < 4; attempt++) { + await consumeCallbackTokenRenewalQuota(env, WS, T0); + } + const nextWindow = T0 + WINDOW_SECONDS * 1000; + + for (let attempt = 0; attempt < 3; attempt++) { + expect(await consumeCallbackTokenRenewalQuota(env, WS, nextWindow)).toEqual({ + outcome: 'allowed', + }); + } + expect((await consumeCallbackTokenRenewalQuota(env, WS, nextWindow)).outcome).toBe('limited'); + expect(counterRows()).toEqual([ + { workspace_id: WS, window_start: T0 / 1000 - 100 + WINDOW_SECONDS, count: 4 }, + ]); + }); + + it('counts each workspace separately', async () => { + const env = makeEnv(); + for (let attempt = 0; attempt < 4; attempt++) { + await consumeCallbackTokenRenewalQuota(env, WS, T0); + } + + expect(await consumeCallbackTokenRenewalQuota(env, OTHER_WS, T0)).toEqual({ + outcome: 'allowed', + }); + }); + + it('loses no increments when attempts arrive together', async () => { + const env = makeEnv(); + + const results = await Promise.all( + Array.from({ length: 6 }, () => consumeCallbackTokenRenewalQuota(env, WS, T0)) + ); + + expect(results.filter((result) => result.outcome === 'allowed')).toHaveLength(3); + expect(results.filter((result) => result.outcome === 'limited')).toHaveLength(3); + expect(counterRows()[0]?.count).toBe(6); + }); + + it('reports a workspace deleted since the caller checked it, without writing a row', async () => { + sqlite.prepare('DELETE FROM workspaces WHERE id = ?').run(WS); + + expect(await consumeCallbackTokenRenewalQuota(makeEnv(), WS, T0)).toEqual({ + outcome: 'workspace_missing', + }); + expect(counterRows()).toEqual([]); + }); +}); + +describe('getCallbackTokenRenewalRateLimit', () => { + it.each([ + [ + 'unset', + {}, + DEFAULT_RATE_LIMITS.CALLBACK_TOKEN_RENEWAL, + DEFAULT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS, + ], + [ + 'configured', + { + RATE_LIMIT_CALLBACK_TOKEN_RENEWAL: '5', + RATE_LIMIT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS: '60', + }, + 5, + 60, + ], + [ + 'invalid', + { + RATE_LIMIT_CALLBACK_TOKEN_RENEWAL: '0', + RATE_LIMIT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS: 'soon', + }, + DEFAULT_RATE_LIMITS.CALLBACK_TOKEN_RENEWAL, + DEFAULT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS, + ], + ])('resolves the limit when %s', (_label, vars, limit, windowSeconds) => { + expect(getCallbackTokenRenewalRateLimit(vars as unknown as Env)).toEqual({ + limit, + windowSeconds, + }); + }); + + it('defaults to 12 attempts per hour', () => { + expect(DEFAULT_RATE_LIMITS.CALLBACK_TOKEN_RENEWAL).toBe(12); + expect(DEFAULT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS).toBe(3600); + }); +}); diff --git a/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts b/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts index a6d3cab94e..ba44b66749 100644 --- a/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts +++ b/apps/api/tests/unit/services/workspace-callback-token-renewal.test.ts @@ -108,7 +108,11 @@ function bindingReadCount(): { count: number; mutateOn: (n: number, sql: string) beforeEach(() => { sqlite = new Database(':memory:'); - createSchemaTables(sqlite, [schema.workspaces, schema.nodes]); + createSchemaTables(sqlite, [ + schema.workspaces, + schema.nodes, + schema.workspaceCallbackTokenRenewalRateLimits, + ]); seed(); onPrepare = null; }); @@ -214,6 +218,48 @@ describe('renewWorkspaceCallbackToken', () => { expect(error.statusCode).toBe(410); }); + it('suppresses the renewed credential when the workspace moves to another node while minting', async () => { + const env = makeEnv(); + const nodeToken = await signNodeCallbackToken(NODE, env); + onPrepare = (sql) => { + if (sql.includes('from "workspaces"') && sql.includes('left join "nodes"')) { + sqlite.prepare('UPDATE workspaces SET node_id = ? WHERE id = ?').run(OTHER_NODE, WS); + } + }; + + const error = await rejection( + renew(await agedToken(WS, 'workspace', 13 * HOUR), env, nodeToken) + ); + expect([error.statusCode, error.error]).toEqual([410, 'GONE']); + expect(sqlite.prepare('SELECT node_id FROM workspaces WHERE id = ?').get(WS)).toEqual({ + node_id: OTHER_NODE, + }); + }); + + it('ends renewal when the workspace is deleted before its attempt is counted', async () => { + const env = makeEnv(); + const nodeToken = await signNodeCallbackToken(NODE, env); + onPrepare = (sql) => { + if (sql.includes('INSERT INTO workspace_callback_token_renewal_rate_limits')) { + sqlite.prepare('DELETE FROM workspaces WHERE id = ?').run(WS); + } + }; + + const error = await rejection( + renew(await agedToken(WS, 'workspace', 13 * HOUR), env, nodeToken) + ); + expect([error.statusCode, error.error]).toEqual([410, 'GONE']); + }); + + it('refuses an Instant (cf-container) workspace, while the VM control renews', async () => { + const aged = await agedToken(WS, 'workspace', 13 * HOUR); + expect((await renew(aged)).renewed).toBe(true); + + sqlite.prepare("UPDATE nodes SET runtime = 'cf-container' WHERE id = ?").run(NODE); + const error = await rejection(renew(aged)); + expect([error.statusCode, error.error]).toEqual([403, 'FORBIDDEN']); + }); + it('keeps the original generation of a legacy token without gen_iat', async () => { const aged = await agedToken(WS, 'workspace', 13 * HOUR); const result = await renew(aged); diff --git a/apps/api/tests/unit/session-sleep-snapshot-wait.test.ts b/apps/api/tests/unit/session-sleep-snapshot-wait.test.ts index c91ef3a221..2cb7523165 100644 --- a/apps/api/tests/unit/session-sleep-snapshot-wait.test.ts +++ b/apps/api/tests/unit/session-sleep-snapshot-wait.test.ts @@ -9,10 +9,12 @@ import { createSchemaTables, createSqliteD1 } from '../helpers/sqlite-d1'; const mocks = vi.hoisted(() => ({ hibernateAgentSessionOnNode: vi.fn(), + hibernateCallbackTokenDelivery: vi.fn(), })); vi.mock('../../src/services/node-agent', () => ({ hibernateAgentSessionOnNode: mocks.hibernateAgentSessionOnNode, + hibernateCallbackTokenDelivery: mocks.hibernateCallbackTokenDelivery, })); describe('waitForFinalSessionSnapshot', () => { @@ -42,7 +44,10 @@ describe('waitForFinalSessionSnapshot', () => { SESSION_SNAPSHOT_POLL_INTERVAL_MS: '1', } as unknown as Env; const db = drizzle(testEnv.DATABASE, { schema }); - mocks.hibernateAgentSessionOnNode.mockResolvedValueOnce({ status: 'pending', accepted: true }); + mocks.hibernateAgentSessionOnNode.mockResolvedValueOnce({ + status: 'pending', + accepted: true, + }); await waitForFinalSessionSnapshot(db, testEnv, { nodeId: 'node-1', @@ -69,11 +74,71 @@ describe('waitForFinalSessionSnapshot', () => { ]); expect(manifest).not.toHaveProperty('acpSessionId'); expect(manifest).not.toHaveProperty('agentType'); - expect(sqlite.prepare(`SELECT capture_error FROM session_snapshots WHERE id = 'snapshot-1'`).get()).toEqual({ + expect( + sqlite.prepare(`SELECT capture_error FROM session_snapshots WHERE id = 'snapshot-1'`).get() + ).toEqual({ capture_error: null, }); } finally { sqlite.close(); } }); + + // The agent installs a delivered workspace token as soon as it reads a hibernate request, + // accepted or not, so polling until it accepts must not sign a new token every poll. + it('gives every poll of one hibernate request the same workspace token delivery', async () => { + const sqlite = new Database(':memory:'); + try { + createSchemaTables(sqlite, [schema.sessionSnapshots]); + sqlite + .prepare( + `INSERT INTO session_snapshots + (id, workspace_id, user_id, chat_session_id, agent_session_id, runtime, status, + degradation, manifest_r2_key, snapshot_generation, capture_generation, capture_error, + expires_at, updated_at) + VALUES ('snapshot-1', 'workspace-1', 'user-1', 'chat-1', 'agent-1', 'vm', + 'available', 'none', 'snapshots/chat-1/previous/manifest.json', 'previous', + 'capture-1', 'capture failed', '2026-08-20T00:00:00.000Z', '2026-08-15T00:00:00.000Z')` + ) + .run(); + const testEnv = { + DATABASE: createSqliteD1(sqlite), + R2: { put: vi.fn(), delete: vi.fn(async () => undefined) }, + SESSION_SNAPSHOT_R2_PREFIX: 'snapshots', + SESSION_SNAPSHOT_REQUEST_TIMEOUT_MS: '5000', + SESSION_SNAPSHOT_PROGRESS_IDLE_TIMEOUT_MS: '1000', + SESSION_SNAPSHOT_POLL_INTERVAL_MS: '1', + } as unknown as Env; + const delivery = vi.fn(async () => 'delivered-token'); + mocks.hibernateCallbackTokenDelivery.mockReset(); + mocks.hibernateCallbackTokenDelivery.mockReturnValueOnce(delivery); + mocks.hibernateAgentSessionOnNode.mockReset(); + mocks.hibernateAgentSessionOnNode + .mockResolvedValueOnce({ status: 'pending', accepted: false }) + .mockResolvedValueOnce({ status: 'pending', accepted: false }) + .mockResolvedValueOnce({ status: 'pending', accepted: true }); + + await waitForFinalSessionSnapshot(drizzle(testEnv.DATABASE, { schema }), testEnv, { + nodeId: 'node-1', + workspaceId: 'workspace-1', + agentSessionId: 'agent-1', + chatSessionId: 'chat-1', + runtime: 'vm', + userId: 'user-1', + }); + + expect(mocks.hibernateCallbackTokenDelivery).toHaveBeenCalledTimes(1); + expect(mocks.hibernateCallbackTokenDelivery).toHaveBeenCalledWith(testEnv, { + workspaceId: 'workspace-1', + nodeId: 'node-1', + }); + const calls = mocks.hibernateAgentSessionOnNode.mock.calls; + expect(calls).toHaveLength(3); + for (const call of calls) { + expect(call[6]).toBe(delivery); + } + } finally { + sqlite.close(); + } + }); }); diff --git a/apps/api/tests/workers/workspace-callback-token-renewal.test.ts b/apps/api/tests/workers/workspace-callback-token-renewal.test.ts index 52aaa99832..12155e14f4 100644 --- a/apps/api/tests/workers/workspace-callback-token-renewal.test.ts +++ b/apps/api/tests/workers/workspace-callback-token-renewal.test.ts @@ -22,8 +22,15 @@ import { signNodeCallbackToken, verifyCallbackToken, } from '../../src/services/jwt'; -import { hibernateAgentSessionOnNode } from '../../src/services/node-agent-session-snapshots'; +import { + hibernateAgentSessionOnNode, + hibernateCallbackTokenDelivery, +} from '../../src/services/node-agent-session-snapshots'; import { mintWorkspaceCallbackTokenForNodeDelivery } from '../../src/services/workspace-callback-token-binding'; +import { + DEFAULT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS, + getCallbackTokenRenewalRateLimit, +} from '../../src/services/workspace-callback-token-renewal-rate-limit'; import { seedNode, seedUser, seedWorkspace } from './helpers/seed-d1'; const testEnv = env as unknown as Env; @@ -49,11 +56,15 @@ const WS_ON_STOPPED_NODE = `${PREFIX}-ws-stopped-node`; const WS_OTHER_OWNER_NODE = `${PREFIX}-ws-other-owner-node`; const WS_MISSING = `${PREFIX}-ws-missing`; const WS_INSTANT = `${PREFIX}-ws-instant`; +const WS_RATE = `${PREFIX}-ws-rate`; +const WS_RATE_CONTROL = `${PREFIX}-ws-rate-control`; +const WS_CASCADE = `${PREFIX}-ws-cascade`; let nodeToken: string; let otherNodeToken: string; let otherOwnerNodeToken: string; let stoppedNodeToken: string; +let instantNodeToken: string; /** * Same claims as production `signCallbackToken` / `signNodeCallbackToken`, with the issue @@ -138,13 +149,25 @@ beforeAll(async () => { // Corrupt binding: placement never puts a workspace on another user's node. await seedWorkspace(WS_OTHER_OWNER_NODE, OTHER_OWNER_NODE_ID, USER_ID, { status: 'running' }); await seedWorkspace(WS_INSTANT, INSTANT_NODE_ID, USER_ID, { status: 'running' }); + await seedWorkspace(WS_RATE, NODE_ID, USER_ID, { status: 'running' }); + await seedWorkspace(WS_RATE_CONTROL, NODE_ID, USER_ID, { status: 'running' }); + await seedWorkspace(WS_CASCADE, NODE_ID, USER_ID, { status: 'running' }); nodeToken = await signNodeCallbackToken(NODE_ID, testEnv); otherNodeToken = await signNodeCallbackToken(OTHER_NODE_ID, testEnv); otherOwnerNodeToken = await signNodeCallbackToken(OTHER_OWNER_NODE_ID, testEnv); stoppedNodeToken = await signNodeCallbackToken(STOPPED_NODE_ID, testEnv); + instantNodeToken = await signNodeCallbackToken(INSTANT_NODE_ID, testEnv); }); +/** Wait out a window boundary that is about to pass, so a burst stays in one window. */ +async function startOfUncrowdedRenewalWindow(): Promise { + const windowSeconds = getCallbackTokenRenewalRateLimit(testEnv).windowSeconds; + const secondsLeft = windowSeconds - (Math.floor(Date.now() / 1000) % windowSeconds); + if (secondsLeft < 15) + await new Promise((resolve) => setTimeout(resolve, (secondsLeft + 1) * 1000)); +} + describe('POST /api/workspaces/:id/callback-token/renew', () => { it('renews an aged token for the hosting node and keeps the generation issue time', async () => { const aged = await signAgedCallbackToken(WS_ACTIVE, 'workspace', 13 * HOUR); @@ -337,6 +360,66 @@ describe('POST /api/workspaces/:id/callback-token/renew', () => { } }); + it('refuses an Instant (cf-container) workspace, which gets a new token when it wakes', async () => { + const aged = await signAgedCallbackToken(WS_INSTANT, 'workspace', 13 * HOUR); + + const response = await renew(WS_INSTANT, aged, { + nodeId: INSTANT_NODE_ID, + nodeToken: instantNodeToken, + }); + + expect(response.status).toBe(403); + expect(await errorCode(response)).toBe('FORBIDDEN'); + }); + + it('limits renewal attempts per workspace, counting only authenticated ones', async () => { + await startOfUncrowdedRenewalWindow(); + const { limit, windowSeconds } = getCallbackTokenRenewalRateLimit(testEnv); + expect([limit, windowSeconds]).toEqual([12, DEFAULT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS]); + const aged = await signAgedCallbackToken(WS_RATE, 'workspace', 13 * HOUR); + + // A caller without the hosting node's credential cannot spend the workspace's attempts. + for (let attempt = 0; attempt < limit; attempt++) { + const rejected = await renew(WS_RATE, aged, { nodeId: NODE_ID, nodeToken: otherNodeToken }); + expect(rejected.status).toBe(403); + } + for (let attempt = 0; attempt < limit; attempt++) { + const allowed = await renew(WS_RATE, aged, { nodeId: NODE_ID, nodeToken }); + expect(allowed.status).toBe(200); + } + + const limited = await renew(WS_RATE, aged, { nodeId: NODE_ID, nodeToken }); + expect(limited.status).toBe(429); + expect(await errorCode(limited)).toBe('RATE_LIMIT_EXCEEDED'); + const retryAfter = Number(limited.headers.get('Retry-After')); + expect(retryAfter).toBeGreaterThanOrEqual(1); + expect(retryAfter).toBeLessThanOrEqual(windowSeconds); + + const otherWorkspace = await renew( + WS_RATE_CONTROL, + await signAgedCallbackToken(WS_RATE_CONTROL, 'workspace', 13 * HOUR), + { nodeId: NODE_ID, nodeToken } + ); + expect(otherWorkspace.status).toBe(200); + expect(((await otherWorkspace.json()) as { renewed: boolean }).renewed).toBe(true); + }); + + it('drops the renewal counter when its workspace row is deleted', async () => { + const aged = await signAgedCallbackToken(WS_CASCADE, 'workspace', 13 * HOUR); + expect((await renew(WS_CASCADE, aged, { nodeId: NODE_ID, nodeToken })).status).toBe(200); + const counted = () => + testEnv.DATABASE.prepare( + 'SELECT COUNT(*) AS n FROM workspace_callback_token_renewal_rate_limits WHERE workspace_id = ?' + ) + .bind(WS_CASCADE) + .first<{ n: number }>('n'); + expect(await counted()).toBe(1); + + await testEnv.DATABASE.prepare('DELETE FROM workspaces WHERE id = ?').bind(WS_CASCADE).run(); + + expect(await counted()).toBe(0); + }); + it('lets a workspace callback that failed after 24h succeed with the renewed token', async () => { const expired = await signAgedCallbackToken(WS_ACTIVE, 'workspace', DAY + HOUR); const rejected = await SELF.fetch( @@ -434,4 +517,41 @@ describe('hibernateAgentSessionOnNode workspace token delivery', () => { expect(body).toEqual({ chatSessionId: 'chat-1', runtime: 'vm', background: true }); }); + + it('mints once for every repeat of a request that shares a delivery', async () => { + const deliver = hibernateCallbackTokenDelivery(testEnv, { + workspaceId: WS_ACTIVE, + nodeId: NODE_ID, + }); + + const first = deliver(); + expect(deliver()).toBe(first); + await expect( + verifyCallbackToken((await first) as string, testEnv, { expectedScope: 'workspace' }) + ).resolves.toMatchObject({ workspace: WS_ACTIVE }); + }); + + it('delivers the token its caller supplies', async () => { + fetchMock.mockResolvedValue( + new Response(JSON.stringify({ status: 'pending', accepted: false }), { status: 202 }) + ); + vi.stubGlobal('fetch', fetchMock); + const deliver = vi.fn(async () => 'token-minted-for-the-first-poll'); + + await hibernateAgentSessionOnNode( + NODE_ID, + WS_ACTIVE, + 'agent-session-1', + testEnv, + USER_ID, + { chatSessionId: 'chat-1', runtime: 'vm', background: true }, + deliver + ); + + expect(deliver).toHaveBeenCalledTimes(1); + const [, init] = fetchMock.mock.calls[0] as [string, RequestInit]; + expect(JSON.parse(String(init.body)).workspaceCallbackToken).toBe( + 'token-minted-for-the-first-poll' + ); + }); }); diff --git a/apps/www/src/content/docs/docs/architecture/security.md b/apps/www/src/content/docs/docs/architecture/security.md index 70eaee272e..246c815aac 100644 --- a/apps/www/src/content/docs/docs/architecture/security.md +++ b/apps/www/src/content/docs/docs/architecture/security.md @@ -107,10 +107,12 @@ VM agents call the API Worker with RS256 callback tokens signed by the Worker (` Both last `CALLBACK_TOKEN_EXPIRY_MS` (24 hours by default). A workspace token is renewed without ever leaving its workspace's scope: -1. **Renewal.** After each successful heartbeat, the VM agent renews workspace tokens that are past `CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO` of their lifetime by calling `POST /api/workspaces/:id/callback-token/renew` with two proofs: the workspace's current, unexpired token and the node's own token. The Worker renews only when D1 binds the workspace to that node, the node belongs to the workspace's owner, and both are still active (`apps/api/src/services/workspace-callback-token-renewal.ts`). A node token alone cannot obtain a workspace token, a workspace token copied out of a devcontainer cannot renew itself, and an expired token is never renewed. -2. **Delivery.** When the Worker asks a VM node to snapshot a session for sleep, the request carries a fresh workspace token over the authenticated node-management channel, the same way workspace creation does. It is minted only if the workspace is still active on that node, and never for Instant containers, which receive a fresh token on every cold wake. +1. **Renewal.** After each successful heartbeat, the VM agent renews workspace tokens that are past `CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO` of their lifetime by calling `POST /api/workspaces/:id/callback-token/renew` with two proofs: the workspace's current, unexpired token and the node's own token. The Worker renews only when D1 binds the workspace to that VM node, the node belongs to the workspace's owner, and both are still active (`apps/api/src/services/workspace-callback-token-renewal.ts`). A node token alone cannot obtain a workspace token, a workspace token copied out of a devcontainer cannot renew itself, and an expired token is never renewed. Authenticated renewal attempts are limited per workspace (`RATE_LIMIT_CALLBACK_TOKEN_RENEWAL`, counted atomically in D1), so a holder of both proofs cannot mint tokens in a loop. +2. **Delivery.** When the Worker asks a VM node to snapshot a session for sleep, the request carries a fresh workspace token over the authenticated node-management channel, the same way workspace creation does. It is minted only if the workspace is still active on that node. -Deleting, stopping or moving a workspace ends renewal, so its callback authority still lapses within one token lifetime. A renewed token keeps its first issue time in a `gen_iat` claim, so the Instant stale-callback guard still recognizes a callback from a replaced container (`apps/api/src/routes/_stale-callback-guard.ts`). +Neither path serves Instant (cf-container) workspaces. Their container receives a fresh token on every cold wake, one per container generation, so a superseded generation cannot extend its authority past the token it started with. + +Deleting or stopping a workspace ends its callback authority at once: every workspace callback checks the workspace status. Moving a workspace to another node ends renewal for the old node, whose existing token then lapses within one token lifetime. A renewed token keeps its first issue time in a `gen_iat` claim, so the Instant stale-callback guard compares generations correctly (`apps/api/src/routes/_stale-callback-guard.ts`). Credentials that an agent process received when it started, such as the SAM AI proxy key, are not rotated inside the running process; the process picks up the current token the next time it starts. From c19872f55f2475685c817874e16413d370316ba6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 09:58:15 +0000 Subject: [PATCH 11/19] fix(vm-agent): lock workspace token reads and follow renewals everywhere Review findings on the agent side: - A new concurrency test found a data race. callbackTokenForWorkspace and workspaceCallbackToken read runtime.CallbackToken through a shared pointer after releasing workspaceMu, while renewal and delivery write it under the lock. Every reader now goes through a locked accessor: SessionHost creation, git credential auth, publish, provisioning and standalone clone (rule 46). - Publish jobs captured the token once for up to DeployBuildPublishTimeout. Their event reporter and control-plane client now read the current workspace token for every request, falling back to the captured one. - A renewed or delivered token whose claims name another workspace, or the node, is never installed. A renewal response like that is retried after backoff; it is not latched. - The refresh-ratio clamp now matches the control plane: non-finite values mean the default, finite values clamp to 0.1-0.9. - New tests: a rate-limited renewal backs off, the reporter built by getOrCreateReporter raises a long pause through the node error reporter, no token value is ever logged, and a SessionHost created during a token change ends with the new token. - Rule 54 now says that a 401 for a replaceable credential pauses and resumes instead of terminating. Co-Authored-By: Claude Opus 5.5 --- .../54-vm-agent-rollout-compatibility.md | 4 +- .../internal/config/callback_token_renewal.go | 21 +- .../config/callback_token_renewal_test.go | 6 +- .../internal/messagereport/credential_test.go | 22 +- .../vm-agent/internal/publish/controlplane.go | 37 +- .../internal/publish/controlplane_test.go | 42 ++ .../internal/server/git_credential.go | 4 +- .../vm-agent/internal/server/mcp_build.go | 31 +- .../internal/server/standalone_workspace.go | 2 +- .../workspace_callback_token_renewal.go | 40 +- .../workspace_callback_token_safety_test.go | 390 ++++++++++++++++++ .../internal/server/workspace_provisioning.go | 38 +- 12 files changed, 581 insertions(+), 56 deletions(-) create mode 100644 packages/vm-agent/internal/server/workspace_callback_token_safety_test.go diff --git a/packages/vm-agent/.claude/rules/54-vm-agent-rollout-compatibility.md b/packages/vm-agent/.claude/rules/54-vm-agent-rollout-compatibility.md index 4f2b761a4c..9807a24552 100644 --- a/packages/vm-agent/.claude/rules/54-vm-agent-rollout-compatibility.md +++ b/packages/vm-agent/.claude/rules/54-vm-agent-rollout-compatibility.md @@ -19,12 +19,12 @@ Required pattern: 10. Missing build metadata is the normal pre-heartbeat state for a freshly booting VM. Cleanup must preserve a configurable boot grace before retiring an unversioned, unclaimed node. 11. A state machine waiting on a claimed node must distinguish missing/deleted state from "still booting" and terminalize promptly without returning the gone node to a reusable pool. 12. Control-plane changes that stop callback storms MUST stand alone for already-deployed agents: terminal statuses and low-severity logging must be correct even if the old VM agent keeps retrying until it is replaced. -13. VM-agent callback loops MUST treat terminal control-plane statuses (`401`, `403`, `404`, `410`) as stop signals, or otherwise use exponential backoff with a hard retry/time budget. Unbounded retries after a terminal resource response are not rollout-compatible. +13. VM-agent callback loops MUST treat terminal control-plane statuses (`401`, `403`, `404`, `410`) as stop signals, or otherwise use exponential backoff with a hard retry/time budget. Unbounded retries after a terminal resource response are not rollout-compatible. A `401` rejects one credential rather than the resource, so a loop whose credential the agent can replace stops presenting the rejected token and resumes when a different token arrives, without discarding queued work: the chat message reporter (`internal/messagereport/credential.go`) and workspace token renewal (`internal/server/workspace_callback_token_renewal.go`) do this. 14. Cloud-init and generated install scripts must request the exact `VM_AGENT_REQUIRED_VERSION` release when it is configured. Unversioned downloads are reserved for legacy/local/manual installs and intentional `skip_agent` deployments. 15. Established deployments must publish Worker revisions in an order that leaves an Instant recovery attempt after the final code update. The Instant recovery budget must enforce the minimum required by that revision count. First-install bootstrap revisions may run only when no prior Worker exists. Tests for scheduling-affecting VM-agent changes should include a stale-but-otherwise-better candidate losing to a compatible node, preferred/warm stale-node rejection, current fresh-node readiness, active stale-node preservation, idle stale-node retirement, and the pre-heartbeat interleaving where a recent bounded warm-node claim exists before any workspace row. -Tests for callback-storm fixes should include old-agent-compatible control-plane assertions for terminal status/severity, plus new-agent assertions that heartbeat, ACP heartbeat, and message-outbox callbacks terminate or exhaust a bounded retry budget after terminal responses. +Tests for callback-storm fixes should include old-agent-compatible control-plane assertions for terminal status/severity, plus new-agent assertions that heartbeat, ACP heartbeat, and message-outbox callbacks terminate or exhaust a bounded retry budget after terminal responses. For the message outbox, `403`/`404`/`410` are terminal, while a `401` must send nothing more on the rejected token and keep the queued rows until a new token arrives. Deployment tests should assert deterministic same-SHA builds, digest-checked reuse without immutable-key overwrites, immutable R2 key construction from the verified VM-agent release SHA, required-version propagation into the download URL, first-install-only bootstrap deployment, and final code publication after the single bulk secret revision. diff --git a/packages/vm-agent/internal/config/callback_token_renewal.go b/packages/vm-agent/internal/config/callback_token_renewal.go index a6d6f46544..080ddf74dd 100644 --- a/packages/vm-agent/internal/config/callback_token_renewal.go +++ b/packages/vm-agent/internal/config/callback_token_renewal.go @@ -1,6 +1,9 @@ package config -import "time" +import ( + "math" + "time" +) // Workspace callback token renewal defaults. Workspace-scoped callback tokens are // minted by the control plane with a fixed lifetime (CALLBACK_TOKEN_EXPIRY_MS, @@ -30,19 +33,15 @@ const ( DefaultWorkspaceCallbackTokenRenewalRetryMax = 30 * time.Minute ) -// clampWorkspaceCallbackTokenRefreshRatio keeps a configured ratio inside the -// supported range; a non-positive or unparseable value falls back to the default. +// clampWorkspaceCallbackTokenRefreshRatio handles a configured ratio exactly as +// the control plane handles CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO +// (shouldRefreshCallbackToken in apps/api/src/services/jwt.ts): a non-finite value +// means the default, and every finite value is clamped to the supported range. func clampWorkspaceCallbackTokenRefreshRatio(ratio float64) float64 { - if ratio <= 0 || ratio != ratio { // ratio != ratio rejects NaN + if math.IsNaN(ratio) || math.IsInf(ratio, 0) { return DefaultWorkspaceCallbackTokenRefreshRatio } - if ratio < MinWorkspaceCallbackTokenRefreshRatio { - return MinWorkspaceCallbackTokenRefreshRatio - } - if ratio > MaxWorkspaceCallbackTokenRefreshRatio { - return MaxWorkspaceCallbackTokenRefreshRatio - } - return ratio + return math.Max(MinWorkspaceCallbackTokenRefreshRatio, math.Min(MaxWorkspaceCallbackTokenRefreshRatio, ratio)) } // positiveDurationOr returns value when positive, else fallback. diff --git a/packages/vm-agent/internal/config/callback_token_renewal_test.go b/packages/vm-agent/internal/config/callback_token_renewal_test.go index 147c843ca3..63a9b6b696 100644 --- a/packages/vm-agent/internal/config/callback_token_renewal_test.go +++ b/packages/vm-agent/internal/config/callback_token_renewal_test.go @@ -45,10 +45,12 @@ func TestWorkspaceCallbackTokenRefreshRatioIsClamped(t *testing.T) { {"0.7", 0.7}, {"0.95", MaxWorkspaceCallbackTokenRefreshRatio}, {"0.01", MinWorkspaceCallbackTokenRefreshRatio}, - {"0", DefaultWorkspaceCallbackTokenRefreshRatio}, - {"-1", DefaultWorkspaceCallbackTokenRefreshRatio}, + // Same handling as the control plane's CALLBACK_TOKEN_REFRESH_THRESHOLD_RATIO. + {"0", MinWorkspaceCallbackTokenRefreshRatio}, + {"-1", MinWorkspaceCallbackTokenRefreshRatio}, {"not-a-number", DefaultWorkspaceCallbackTokenRefreshRatio}, {"NaN", DefaultWorkspaceCallbackTokenRefreshRatio}, + {"Inf", DefaultWorkspaceCallbackTokenRefreshRatio}, } { t.Run(tc.value, func(t *testing.T) { cfg := loadRenewalConfig(t, map[string]string{"WORKSPACE_CALLBACK_TOKEN_REFRESH_RATIO": tc.value}) diff --git a/packages/vm-agent/internal/messagereport/credential_test.go b/packages/vm-agent/internal/messagereport/credential_test.go index 18bbffc91e..0d137618ef 100644 --- a/packages/vm-agent/internal/messagereport/credential_test.go +++ b/packages/vm-agent/internal/messagereport/credential_test.go @@ -6,6 +6,7 @@ import ( "net/http/httptest" "strings" "sync" + "sync/atomic" "testing" "time" ) @@ -97,6 +98,18 @@ func (cp *tokenGatedControlPlane) persistedIDs() map[string]int { return out } +// setTestClock installs a clock the test advances. The reporter reads now under +// r.mu, so installing it under r.mu and advancing it atomically stays race-free +// even if a background flush runs while the test moves the clock. +func setTestClock(r *Reporter, start time.Time) (advance func(time.Duration)) { + var nanos atomic.Int64 + nanos.Store(start.UnixNano()) + r.mu.Lock() + r.now = func() time.Time { return time.Unix(0, nanos.Load()).UTC() } + r.mu.Unlock() + return func(d time.Duration) { nanos.Add(int64(d)) } +} + // newHeldTestReporter builds a reporter whose background loop never ticks during // the test, so every flush below is one the test drives. func newHeldTestReporter(t *testing.T, endpoint, token string, adjust func(*Config)) (*Reporter, func() int) { @@ -247,19 +260,18 @@ func TestCredentialRejection_SurfacesLongPauseOnceWithoutDeleting(t *testing.T) reports = append(reports, info) } }) - clock := time.Date(2026, 10, 4, 8, 0, 0, 0, time.UTC) - r.now = func() time.Time { return clock } + advance := setTestClock(r, time.Date(2026, 10, 4, 8, 0, 0, 0, time.UTC)) enqueueAssistant(t, r, "m0", "m1") r.flush() // 401: pause starts - clock = clock.Add(30 * time.Minute) + advance(30 * time.Minute) r.flush() if len(reports) != 0 { t.Fatalf("pause reported before the budget: %+v", reports) } - clock = clock.Add(31 * time.Minute) + advance(31 * time.Minute) r.flush() r.flush() if len(reports) != 1 { @@ -284,7 +296,7 @@ func TestCredentialRejection_SurfacesLongPauseOnceWithoutDeleting(t *testing.T) r.SetToken("rejected-again") enqueueAssistant(t, r, "m2") r.flush() - clock = clock.Add(2 * time.Hour) + advance(2 * time.Hour) r.flush() if len(reports) != 2 { t.Fatalf("a new pause must be reported again, got %d reports", len(reports)) diff --git a/packages/vm-agent/internal/publish/controlplane.go b/packages/vm-agent/internal/publish/controlplane.go index d07d017451..8ab9f09be2 100644 --- a/packages/vm-agent/internal/publish/controlplane.go +++ b/packages/vm-agent/internal/publish/controlplane.go @@ -27,10 +27,11 @@ const maxControlPlaneErrorBodyBytes = 4096 // HTTPControlPlane talks to the SAM control plane over HTTP using the workspace // callback JWT. It is the production ControlPlane implementation. type HTTPControlPlane struct { - baseURL string - token string - client *http.Client - log *slog.Logger + baseURL string + token string + tokenSource func() string + client *http.Client + log *slog.Logger } // HTTPControlPlaneOptions configures a new HTTPControlPlane. BaseURL, Token, and @@ -38,8 +39,12 @@ type HTTPControlPlane struct { type HTTPControlPlaneOptions struct { BaseURL string Token string - Client *http.Client - Logger *slog.Logger + // TokenSource, when set, supplies the callback JWT for each request, so a + // publish that outlives a workspace token renewal uses the renewed token. + // Token is used whenever it returns "". + TokenSource func() string + Client *http.Client + Logger *slog.Logger } // NewHTTPControlPlane constructs an HTTPControlPlane. @@ -53,13 +58,23 @@ func NewHTTPControlPlane(opts HTTPControlPlaneOptions) *HTTPControlPlane { client = http.DefaultClient } return &HTTPControlPlane{ - baseURL: strings.TrimRight(opts.BaseURL, "/"), - token: opts.Token, - client: client, - log: log.With("component", "publish-controlplane"), + baseURL: strings.TrimRight(opts.BaseURL, "/"), + token: opts.Token, + tokenSource: opts.TokenSource, + client: client, + log: log.With("component", "publish-controlplane"), } } +func (c *HTTPControlPlane) currentToken() string { + if c.tokenSource != nil { + if token := strings.TrimSpace(c.tokenSource()); token != "" { + return token + } + } + return c.token +} + func (c *HTTPControlPlane) InitArtifactUploads(ctx context.Context, projectID string, req ArtifactUploadRequest) (*ArtifactUploadInitResponse, error) { var result ArtifactUploadInitResponse if err := c.do(ctx, projectID, routeComposeImageArtifactsInit, req, &result); err != nil { @@ -111,7 +126,7 @@ func (c *HTTPControlPlane) do(ctx context.Context, projectID, route string, reqB if err != nil { return fmt.Errorf("create request: %w", err) } - req.Header.Set("Authorization", "Bearer "+c.token) + req.Header.Set("Authorization", "Bearer "+c.currentToken()) req.Header.Set("Content-Type", "application/json") resp, err := c.client.Do(req) diff --git a/packages/vm-agent/internal/publish/controlplane_test.go b/packages/vm-agent/internal/publish/controlplane_test.go index e364f9641e..a8e76aef6c 100644 --- a/packages/vm-agent/internal/publish/controlplane_test.go +++ b/packages/vm-agent/internal/publish/controlplane_test.go @@ -169,3 +169,45 @@ func TestHTTPControlPlaneSubmitReleaseSendsReleasePayload(t *testing.T) { t.Fatalf("result = %+v", result) } } + +// A publish can outlive a workspace callback token renewal, so each request reads +// the current token; the token captured at start is the fallback. +func TestHTTPControlPlaneUsesTheCurrentTokenForEachRequest(t *testing.T) { + var gotAuth []string + srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + gotAuth = append(gotAuth, r.Header.Get("Authorization")) + w.Header().Set("Content-Type", "application/json") + _, _ = w.Write([]byte(`{"ok":true}`)) + })) + defer srv.Close() + + current := "token-at-start" + cp := NewHTTPControlPlane(HTTPControlPlaneOptions{ + BaseURL: srv.URL, + Token: "token-captured-at-start", + TokenSource: func() string { return current }, + Client: srv.Client(), + }) + complete := func() { + t.Helper() + if err := cp.CompleteArtifactUploads(context.Background(), "proj1", ArtifactCompleteRequest{}); err != nil { + t.Fatalf("CompleteArtifactUploads: %v", err) + } + } + + complete() + current = "token-renewed-mid-publish" + complete() + current = "" + complete() + + want := []string{"Bearer token-at-start", "Bearer token-renewed-mid-publish", "Bearer token-captured-at-start"} + if len(gotAuth) != len(want) { + t.Fatalf("requests = %v, want %v", gotAuth, want) + } + for i := range want { + if gotAuth[i] != want[i] { + t.Fatalf("request %d authorization = %q, want %q", i, gotAuth[i], want[i]) + } + } +} diff --git a/packages/vm-agent/internal/server/git_credential.go b/packages/vm-agent/internal/server/git_credential.go index a3377fb75c..79ff0dd3ea 100644 --- a/packages/vm-agent/internal/server/git_credential.go +++ b/packages/vm-agent/internal/server/git_credential.go @@ -359,10 +359,10 @@ func (s *Server) callbackAuthCandidates(workspaceID string) []callbackAuthCandid if workspaceID == "" { return candidates } - if runtime, ok := s.getWorkspaceRuntime(workspaceID); ok { + if token, ok := s.lookupWorkspaceCallbackToken(workspaceID); ok { candidates = append(candidates, callbackAuthCandidate{ source: "workspace", - token: strings.TrimSpace(runtime.CallbackToken), + token: token, }) } return candidates diff --git a/packages/vm-agent/internal/server/mcp_build.go b/packages/vm-agent/internal/server/mcp_build.go index 4c6e52be33..ff93524c12 100644 --- a/packages/vm-agent/internal/server/mcp_build.go +++ b/packages/vm-agent/internal/server/mcp_build.go @@ -172,7 +172,7 @@ func (s *Server) handleMcpBuildAndPublishJobStart(w http.ResponseWriter, r *http s.publishJobsMu.Unlock() s.persistVMJobStart(jobID, vmJobKindPublish, prepared.WorkspaceID, vmJobStatusStarting, "starting") - controlPlaneReporter := newPublishJobReporter(s.config.ControlPlaneURL, prepared.ProjectID, jobID, prepared.Token, s.controlPlaneHTTPClient(publishTimeout), prepared.Log) + controlPlaneReporter := newPublishJobReporter(s.config.ControlPlaneURL, prepared.ProjectID, jobID, s.publishCallbackToken(prepared), s.controlPlaneHTTPClient(publishTimeout), prepared.Log) reporter := publish.EventFunc(func(ctx context.Context, event publish.Event) { s.persistPublishEvent(jobID, event) controlPlaneReporter.Event(ctx, event) @@ -236,7 +236,7 @@ func (s *Server) prepareMcpBuildAndPublish(w http.ResponseWriter, r *http.Reques return nil, false } - token := strings.TrimSpace(runtime.CallbackToken) + token := s.workspaceCallbackToken(workspaceID) if token == "" { writeError(w, http.StatusInternalServerError, "workspace has no callback token for publishing") return nil, false @@ -340,10 +340,11 @@ func (s *Server) runPreparedBuildAndPublish(ctx context.Context, prepared *prepa orch := publish.New(publish.Options{ ControlPlane: publish.NewHTTPControlPlane(publish.HTTPControlPlaneOptions{ - BaseURL: s.config.ControlPlaneURL, - Token: prepared.Token, - Client: s.controlPlaneHTTPClient(s.deployBuildPublishTimeout()), - Logger: log, + BaseURL: s.config.ControlPlaneURL, + Token: prepared.Token, + TokenSource: s.publishCallbackToken(prepared), + Client: s.controlPlaneHTTPClient(s.deployBuildPublishTimeout()), + Logger: log, }), Docker: publish.NewHostDocker(), Events: events, @@ -363,16 +364,28 @@ func (s *Server) runPreparedBuildAndPublish(ctx context.Context, prepared *prepa return result, nil } +// publishCallbackToken reads the workspace token for each publish callback, so a +// job that runs while the token is renewed uses the renewed token. The token +// captured when the job was accepted is the fallback if the runtime is gone. +func (s *Server) publishCallbackToken(prepared *preparedBuildPublish) func() string { + return func() string { + if token := s.workspaceCallbackToken(prepared.WorkspaceID); token != "" { + return token + } + return prepared.Token + } +} + type publishJobReporter struct { baseURL string projectID string jobID string - token string + token func() string client *http.Client log *slog.Logger } -func newPublishJobReporter(baseURL, projectID, jobID, token string, client *http.Client, log *slog.Logger) *publishJobReporter { +func newPublishJobReporter(baseURL, projectID, jobID string, token func() string, client *http.Client, log *slog.Logger) *publishJobReporter { return &publishJobReporter{ baseURL: strings.TrimRight(baseURL, "/"), projectID: projectID, @@ -401,7 +414,7 @@ func (r *publishJobReporter) Event(ctx context.Context, event publish.Event) { r.log.Warn("create publish job event request failed", "error", err) return } - req.Header.Set("Authorization", "Bearer "+r.token) + req.Header.Set("Authorization", "Bearer "+r.token()) req.Header.Set("Content-Type", "application/json") resp, err := r.client.Do(req) if err != nil { diff --git a/packages/vm-agent/internal/server/standalone_workspace.go b/packages/vm-agent/internal/server/standalone_workspace.go index 91bdc45b2c..20ac35721c 100644 --- a/packages/vm-agent/internal/server/standalone_workspace.go +++ b/packages/vm-agent/internal/server/standalone_workspace.go @@ -216,7 +216,7 @@ func (s *Server) standaloneCloneSpec(ctx context.Context, runtime *WorkspaceRunt repositoryURL := gitrepo.NormalizeURL(runtime.Repository) var tokenResponse *gitTokenResponse - if callbackToken := strings.TrimSpace(runtime.CallbackToken); callbackToken != "" { + if callbackToken := s.runtimeCallbackToken(runtime); callbackToken != "" { resp, err := s.fetchGitTokenResponseForWorkspace(ctx, runtime.ID, callbackToken) if err != nil { slog.Warn("Standalone repository clone proceeding without git token", "workspace", runtime.ID, "error", err) diff --git a/packages/vm-agent/internal/server/workspace_callback_token_renewal.go b/packages/vm-agent/internal/server/workspace_callback_token_renewal.go index 8b3c2c6b3f..bceb6177c2 100644 --- a/packages/vm-agent/internal/server/workspace_callback_token_renewal.go +++ b/packages/vm-agent/internal/server/workspace_callback_token_renewal.go @@ -169,6 +169,25 @@ func callbackTokenRenewalDue(token string, now time.Time, ratio float64) bool { return !now.Before(issuedAt.Add(time.Duration(float64(lifetime) * ratio))) } +// callbackTokenNamesOtherWorkspace reports whether a token's own claims say it is +// not workspaceID's workspace token: node-scoped, or bound to another workspace +// (internal/auth/jwt.go applies the same rule to tokens it verifies). Claims are +// read WITHOUT verifying the signature; this only keeps a control-plane mistake +// from being installed. A token whose claims cannot be read is left for the +// control plane to reject. +func callbackTokenNamesOtherWorkspace(token, workspaceID string) bool { + var claims struct { + jwt.RegisteredClaims + Workspace string `json:"workspace"` + Scope string `json:"scope"` + } + if _, _, err := jwt.NewParser().ParseUnverified(token, &claims); err != nil { + return false + } + return claims.Scope == "node" || claims.Workspace != workspaceID || + (claims.Subject != "" && claims.Subject != workspaceID) +} + // callbackTokenLifetime reads iat/exp WITHOUT verifying the signature. It only // schedules renewal; the control plane verifies every token it receives. func callbackTokenLifetime(token string) (issuedAt, expiresAt time.Time, ok bool) { @@ -292,7 +311,12 @@ func (s *Server) requestWorkspaceCallbackTokenRenewal( if parsed.Error != "" { detail += " " + parsed.Error } - return classifyWorkspaceTokenRenewal(resp.StatusCode, parsed.Renewed, parsed.Token, parsed.Error, parseErr), strings.TrimSpace(parsed.Token), detail + outcome = classifyWorkspaceTokenRenewal(resp.StatusCode, parsed.Renewed, parsed.Token, parsed.Error, parseErr) + renewed = strings.TrimSpace(parsed.Token) + if outcome == renewalSucceeded && callbackTokenNamesOtherWorkspace(renewed, candidate.workspaceID) { + return renewalTransientFailure, "", detail + " (renewed token is not this workspace's token)" + } + return outcome, renewed, detail } func classifyWorkspaceTokenRenewal(status int, renewed bool, token, errorCode string, parseErr error) workspaceTokenRenewalOutcome { @@ -341,16 +365,22 @@ func (s *Server) replaceRenewedWorkspaceCallbackToken(workspaceID, renewedFrom, } // adoptWorkspaceCallbackTokenLocked installs a token delivered for an existing -// workspace (create, restore, hibernate) unless it would replace the current -// token with one that expires earlier, so a reordered or late delivery never -// rolls a workspace back to an older credential. Caller holds workspaceMu and -// must persist and then propagate when it returns true. +// workspace (create, restore, hibernate) unless its claims name another +// workspace, or it would replace the current token with one that expires +// earlier, so a reordered or late delivery never rolls a workspace back to an +// older credential. Caller holds workspaceMu and must persist and then +// propagate when it returns true. func adoptWorkspaceCallbackTokenLocked(runtime *WorkspaceRuntime, delivered string) bool { delivered = strings.TrimSpace(delivered) current := strings.TrimSpace(runtime.CallbackToken) if delivered == "" || delivered == current { return false } + if callbackTokenNamesOtherWorkspace(delivered, runtime.ID) { + slog.Warn("Ignoring delivered callback token that is not this workspace's token", + "workspace", runtime.ID) + return false + } if current != "" { _, deliveredExpiry, deliveredOK := callbackTokenLifetime(delivered) _, currentExpiry, currentOK := callbackTokenLifetime(current) diff --git a/packages/vm-agent/internal/server/workspace_callback_token_safety_test.go b/packages/vm-agent/internal/server/workspace_callback_token_safety_test.go new file mode 100644 index 0000000000..0cf9b97196 --- /dev/null +++ b/packages/vm-agent/internal/server/workspace_callback_token_safety_test.go @@ -0,0 +1,390 @@ +package server + +// Safety properties of workspace callback token renewal and delivery that sit +// beside the core lifecycle in workspace_callback_token_renewal_test.go: tokens +// that are not this workspace's are never installed, a rate-limited renewal +// backs off instead of latching, a long delivery pause reaches the node's error +// channel, no token value is ever logged, and a SessionHost created while a +// token changes still ends up with the new token. + +import ( + "bytes" + "context" + "io" + "log/slog" + "net/http" + "net/http/httptest" + "path/filepath" + "strings" + "sync" + "testing" + "time" + + "github.com/golang-jwt/jwt/v5" + + "github.com/workspace/vm-agent/internal/acp" + "github.com/workspace/vm-agent/internal/agentsessions" + "github.com/workspace/vm-agent/internal/config" + "github.com/workspace/vm-agent/internal/errorreport" + "github.com/workspace/vm-agent/internal/messagereport" + "github.com/workspace/vm-agent/internal/persistence" + "github.com/workspace/vm-agent/internal/publish" +) + +func nodeScopedTestToken(t *testing.T, nodeID string, issuedAt time.Time) string { + t.Helper() + token, err := jwt.NewWithClaims(jwt.SigningMethodHS256, jwt.MapClaims{ + "workspace": nodeID, + "type": "callback", + "scope": "node", + "sub": nodeID, + "aud": "workspace-callback", + "iat": issuedAt.Unix(), + "exp": issuedAt.Add(24 * time.Hour).Unix(), + }).SignedString([]byte("test-signing-key")) + if err != nil { + t.Fatal(err) + } + return token +} + +func TestWorkspaceTokenRenewal_RateLimitedRenewalBacksOffInsteadOfLatching(t *testing.T) { + h := newRenewalHarness(t, renewalEpoch) + h.addWorkspace(renewalTestWorkspace, workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour)) + h.cp.respondWith(http.StatusTooManyRequests, `{"error":"RATE_LIMIT_EXCEEDED","message":"Too many requests. Please try again later."}`) + + h.pass(13 * time.Hour) + h.pass(30 * time.Second) + if got := len(h.cp.renewalRequests()); got != 1 { + t.Fatalf("retried a rate-limited renewal before the backoff: %d requests", got) + } + + renewed := workspaceTestToken(t, renewalTestWorkspace, h.clock, 24*time.Hour) + h.cp.renewWith(renewed) + h.pass(30 * time.Second) + if h.token(renewalTestWorkspace) != renewed { + t.Fatal("a rate-limited renewal latched the token instead of retrying after the backoff") + } +} + +func TestWorkspaceTokenRenewal_IgnoresARenewedTokenForAnotherWorkspace(t *testing.T) { + h := newRenewalHarness(t, renewalEpoch) + current := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour) + h.addWorkspace(renewalTestWorkspace, current) + foreign := workspaceTestToken(t, "ws-2", renewalEpoch.Add(13*time.Hour), 24*time.Hour) + h.cp.respondWith(http.StatusOK, `{"renewed":true,"token":"`+foreign+`"}`) + + h.pass(13 * time.Hour) + if h.token(renewalTestWorkspace) != current { + t.Fatal("installed a renewed token that names another workspace") + } + + // Treated as a transient control-plane fault: retried after the backoff, not latched. + renewed := workspaceTestToken(t, renewalTestWorkspace, h.clock, 24*time.Hour) + h.cp.renewWith(renewed) + h.pass(time.Minute) + if h.token(renewalTestWorkspace) != renewed { + t.Fatal("renewal did not recover with this workspace's token") + } +} + +func TestWorkspaceTokenDelivery_IgnoresTokensThatAreNotThisWorkspaces(t *testing.T) { + h := newRenewalHarness(t, renewalEpoch) + current := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour) + h.addWorkspace(renewalTestWorkspace, current) + host := acp.NewSessionHost(acp.SessionHostConfig{GatewayConfig: acp.GatewayConfig{CallbackToken: current}}) + t.Cleanup(host.Stop) + h.s.sessionHosts[renewalTestWorkspace+":agent-1"] = host + + later := renewalEpoch.Add(time.Hour) + for name, token := range map[string]string{ + "another workspace's token": workspaceTestToken(t, "ws-2", later, 24*time.Hour), + "the node's own token": nodeScopedTestToken(t, renewalTestNode, later), + } { + h.s.upsertWorkspaceRuntime(renewalTestWorkspace, "", "", "", token) + if h.token(renewalTestWorkspace) != current || !host.UsesCallbackToken(current) { + t.Fatalf("adopted %s as the workspace token", name) + } + } + + own := workspaceTestToken(t, renewalTestWorkspace, later, 24*time.Hour) + h.s.upsertWorkspaceRuntime(renewalTestWorkspace, "", "", "", own) + if h.token(renewalTestWorkspace) != own || !host.UsesCallbackToken(own) { + t.Fatal("did not adopt this workspace's own newer token") + } +} + +// The reporter getOrCreateReporter builds must raise the long-pause alert through +// the node's error reporter, which authenticates with the node token and so still +// gets through while the workspace token is refused. +func TestGetOrCreateReporter_RaisesALongMessagePauseThroughTheNodeErrorReporter(t *testing.T) { + t.Setenv("MSG_AUTH_RENEWAL_WAIT", "1ms") + t.Setenv("MSG_BATCH_MAX_WAIT", "10ms") + const rejected = "workspace-token-the-control-plane-refuses" + + var mu sync.Mutex + var errorBodies []string + messageAttempts := 0 + controlPlane := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + body, _ := io.ReadAll(r.Body) + switch { + case r.Method == http.MethodPost && r.URL.Path == "/api/nodes/node-1/errors": + mu.Lock() + errorBodies = append(errorBodies, string(body)) + mu.Unlock() + w.WriteHeader(http.StatusNoContent) + case r.Method == http.MethodPost && r.URL.Path == "/api/workspaces/ws-1/messages": + mu.Lock() + messageAttempts++ + mu.Unlock() + writeRenewalJSON(w, http.StatusUnauthorized, `{"error":"UNAUTHORIZED","message":"Invalid or expired callback token"}`) + default: + http.NotFound(w, r) + } + })) + t.Cleanup(controlPlane.Close) + + dir := t.TempDir() + errorReporter := errorreport.New(controlPlane.URL, "node-1", "node-token", errorreport.Config{ + DBPath: filepath.Join(dir, "error-reports.db"), + FlushInterval: 10 * time.Millisecond, + }) + errorReporter.Start() + t.Cleanup(errorReporter.Shutdown) + + s := &Server{ + config: &config.Config{ + NodeID: "node-1", + ControlPlaneURL: controlPlane.URL, + PersistenceDBPath: filepath.Join(dir, "state.db"), + }, + errorReporter: errorReporter, + messageReporters: map[string]*messagereport.Reporter{}, + workspaces: map[string]*WorkspaceRuntime{ + "ws-1": {ID: "ws-1", CallbackToken: rejected, Status: "running"}, + }, + } + reporter := s.getOrCreateReporter("ws-1", "proj-1", "chat-1") + if reporter == nil { + t.Fatal("getOrCreateReporter returned nil") + } + t.Cleanup(reporter.Shutdown) + if err := reporter.Enqueue(messagereport.Message{MessageID: "m-1", Role: "assistant", Content: "reply"}); err != nil { + t.Fatal(err) + } + + waitFor(t, func() bool { + mu.Lock() + defer mu.Unlock() + for _, body := range errorBodies { + if strings.Contains(body, "messagereport.credential_wait") { + return true + } + } + return false + }, "the node error reporter never received the message-persistence pause") + + mu.Lock() + defer mu.Unlock() + if messageAttempts == 0 { + t.Fatal("the reporter never tried to deliver, so the pause was not caused by a refusal") + } + for _, body := range errorBodies { + if strings.Contains(body, rejected) { + t.Fatal("the error report carried the workspace token") + } + } +} + +type lockedLogBuffer struct { + mu sync.Mutex + buf bytes.Buffer +} + +func (b *lockedLogBuffer) Write(p []byte) (int, error) { + b.mu.Lock() + defer b.mu.Unlock() + return b.buf.Write(p) +} + +func (b *lockedLogBuffer) String() string { + b.mu.Lock() + defer b.mu.Unlock() + return b.buf.String() +} + +// Every renewal outcome and every refusal path logs workspace IDs, statuses and +// error codes, never a token. +func TestWorkspaceTokenRenewal_NeverLogsATokenValue(t *testing.T) { + logs := &lockedLogBuffer{} + previous := slog.Default() + slog.SetDefault(slog.New(slog.NewTextHandler(logs, &slog.HandlerOptions{Level: slog.LevelDebug}))) + t.Cleanup(func() { slog.SetDefault(previous) }) + + const nodeToken = "node-token-that-must-never-be-logged" + h := newRenewalHarness(t, renewalEpoch) + h.s.callbackToken = nodeToken + h.s.config.CallbackToken = nodeToken + h.cp.nodeToken = nodeToken + first := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour) + h.addWorkspace(renewalTestWorkspace, first) + tokens := []string{nodeToken, first} + + h.cp.respondWith(http.StatusServiceUnavailable, `{"error":"SERVICE_UNAVAILABLE"}`) + h.pass(13 * time.Hour) // transient failure + foreign := workspaceTestToken(t, "ws-2", h.clock, 24*time.Hour) + tokens = append(tokens, foreign) + h.cp.respondWith(http.StatusOK, `{"renewed":true,"token":"`+foreign+`"}`) + h.pass(time.Hour) // a renewed token for another workspace + renewed := workspaceTestToken(t, renewalTestWorkspace, h.clock, 24*time.Hour) + tokens = append(tokens, renewed) + h.cp.renewWith(renewed) + h.pass(time.Hour) // success + h.s.upsertWorkspaceRuntime(renewalTestWorkspace, "", "", "", foreign) // refused delivery + h.cp.respondWith(http.StatusUnauthorized, `{"error":"UNAUTHORIZED","message":"Invalid or expired callback token"}`) + h.pass(13 * time.Hour) // refusal + + _, db := openTestSQLiteDB(t) + reporter, err := messagereport.New(db, messagereport.Config{ + BatchMaxWait: 10 * time.Millisecond, Endpoint: h.cp.server.URL, + WorkspaceID: renewalTestWorkspace, ProjectID: "proj-1", SessionID: "chat-1", + }) + if err != nil { + t.Fatal(err) + } + t.Cleanup(reporter.Shutdown) + reporter.SetToken(renewed) // the fake control plane only accepts tokens it renewed + h.cp.mu.Lock() + delete(h.cp.accepted, renewed) + h.cp.mu.Unlock() + if err := reporter.Enqueue(messagereport.Message{MessageID: "m-1", Role: "assistant", Content: "reply"}); err != nil { + t.Fatal(err) + } + waitFor(t, func() bool { return heldOnToken(h.cp, renewed) }, "the reporter never hit a 401") + + output := logs.String() + for _, want := range []string{ + "Workspace callback token renewal failed", + "Workspace callback token renewed", + "Ignoring delivered callback token that is not this workspace's token", + "Control plane refused workspace callback token renewal", + "holding messages until it is replaced", + } { + if !strings.Contains(output, want) { + t.Fatalf("expected log %q was not captured, so the absence check below proves nothing:\n%s", want, output) + } + } + for _, token := range tokens { + signature := token[strings.LastIndex(token, ".")+1:] + if strings.Contains(output, token) || strings.Contains(output, signature) { + t.Fatalf("a token value was logged:\n%s", output) + } + } +} + +// A SessionHost created while a renewal installs a new token must end up with the +// new token whichever runs first: creation reads the token and registers the host +// under one sessionHostMu hold, and propagation updates hosts under the same lock. +func TestWorkspaceTokenRenewal_HostCreatedDuringATokenChangeGetsTheNewToken(t *testing.T) { + for round := 0; round < 200; round++ { + h := newRenewalHarness(t, renewalEpoch) + h.s.config.ACPMessageBufferSize = 8 + h.s.config.ACPViewerSendBuffer = 2 + h.s.agentSessions = agentsessions.NewManager() + h.s.sessionMcpServers = map[string][]acp.McpServerEntry{} + h.s.sessionProfileOvr = map[string]profileOverrides{} + h.s.sessionTaskCtx = map[string]taskCallbackContext{} + old := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch, 24*time.Hour) + h.addWorkspace(renewalTestWorkspace, old) + renewed := workspaceTestToken(t, renewalTestWorkspace, renewalEpoch.Add(time.Duration(round+1)*time.Minute), 24*time.Hour) + + var host *acp.SessionHost + var wg sync.WaitGroup + start := make(chan struct{}) + wg.Add(2) + go func() { + defer wg.Done() + <-start + host = h.s.getOrCreateSessionHost(renewalTestWorkspace+":agent-1", renewalTestWorkspace, "agent-1", + agentsessions.Session{ID: "agent-1", WorkspaceID: renewalTestWorkspace}, nil, "") + }() + go func() { + defer wg.Done() + <-start + h.s.replaceRenewedWorkspaceCallbackToken(renewalTestWorkspace, old, renewed) + }() + close(start) + wg.Wait() + + if host == nil { + t.Fatal("no SessionHost was created") + } + if !host.UsesCallbackToken(renewed) { + host.Stop() + t.Fatalf("round %d: a SessionHost created during the token change kept the old token", round) + } + host.Stop() + } +} + +// A publish job runs for up to DeployBuildPublishTimeout. Its callbacks must use a +// workspace token renewed while it runs, not the one captured when it started. +func TestPublishJob_CallbacksUseATokenRenewedWhileTheJobRuns(t *testing.T) { + s, key := mcpBuildTestServer(t) + tmp := t.TempDir() + t.Setenv("SAM_DOCKER_CLI_PATH", fakeDockerCLI(t, tmp, "", true)) + store, err := persistence.Open(filepath.Join(tmp, "vm-agent.db")) + if err != nil { + t.Fatal(err) + } + t.Cleanup(func() { _ = store.Close() }) + s.store = store + + var mu sync.Mutex + var eventAuth []string + callbacks := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + if strings.HasSuffix(r.URL.Path, "/deployment-publish-jobs/job-renew/events") { + mu.Lock() + eventAuth = append(eventAuth, r.Header.Get("Authorization")) + mu.Unlock() + } + w.Header().Set("Content-Type", "application/json") + _, _ = w.Write([]byte(`{"ok":true}`)) + })) + t.Cleanup(callbacks.Close) + s.config.ControlPlaneURL = callbacks.URL + s.workspaces["ws-001"] = &WorkspaceRuntime{ + ID: "ws-001", Status: "running", WorkspaceDir: "/workspace/WS_001", ProjectID: "proj-1", + CallbackToken: "token-when-the-job-started", + } + + started := make(chan struct{}) + release := make(chan struct{}) + s.buildPublishRunner = func(context.Context, *preparedBuildPublish, publish.EventSink) (*publish.ReleaseResult, error) { + close(started) + <-release + return &publish.ReleaseResult{ReleaseID: "rel-1", Version: 1, Status: "created"}, nil + } + rec := mcpBuildJobStartPOST(t, s, key, "ws-001", "job-renew", McpBuildAndPublishRequest{ + PublishJobID: "job-renew", Environment: "staging", EnvironmentID: "env-1", + }, context.Background()) + if rec.Code != http.StatusAccepted { + t.Fatalf("expected 202, got %d: %s", rec.Code, rec.Body.String()) + } + <-started + if !s.replaceRenewedWorkspaceCallbackToken("ws-001", "token-when-the-job-started", "token-renewed-during-the-job") { + t.Fatal("renewal was not installed") + } + close(release) + + waitFor(t, func() bool { + mu.Lock() + defer mu.Unlock() + return len(eventAuth) > 0 && eventAuth[len(eventAuth)-1] == "Bearer token-renewed-during-the-job" + }, "the publish job's later callbacks did not use the renewed token") + mu.Lock() + defer mu.Unlock() + if eventAuth[0] != "Bearer token-when-the-job-started" { + t.Fatalf("first publish callback used %q, want the token the job started with", eventAuth[0]) + } +} diff --git a/packages/vm-agent/internal/server/workspace_provisioning.go b/packages/vm-agent/internal/server/workspace_provisioning.go index 45118a97fc..abb5506742 100644 --- a/packages/vm-agent/internal/server/workspace_provisioning.go +++ b/packages/vm-agent/internal/server/workspace_provisioning.go @@ -26,20 +26,42 @@ type workspaceRuntimeMetadataResponse struct { } func (s *Server) callbackTokenForWorkspace(workspaceID string) string { - if runtime, ok := s.getWorkspaceRuntime(workspaceID); ok { - if token := strings.TrimSpace(runtime.CallbackToken); token != "" { - return token - } + if token := s.workspaceCallbackToken(workspaceID); token != "" { + return token } return strings.TrimSpace(s.config.CallbackToken) } +// Renewal and control-plane delivery replace runtime.CallbackToken while the +// workspace runs (workspace_callback_token_renewal.go), always under +// workspaceMu, so every read takes the lock too (rule 46). Callers must not hold +// workspaceMu. + func (s *Server) workspaceCallbackToken(workspaceID string) string { - if runtime, ok := s.getWorkspaceRuntime(workspaceID); ok { - return strings.TrimSpace(runtime.CallbackToken) + token, _ := s.lookupWorkspaceCallbackToken(workspaceID) + return token +} + +// lookupWorkspaceCallbackToken also reports whether the workspace exists. +func (s *Server) lookupWorkspaceCallbackToken(workspaceID string) (string, bool) { + s.workspaceMu.RLock() + defer s.workspaceMu.RUnlock() + runtime, ok := s.workspaces[workspaceID] + if !ok || runtime == nil { + return "", false + } + return strings.TrimSpace(runtime.CallbackToken), true +} + +// runtimeCallbackToken reads the token of a runtime the caller already holds. +func (s *Server) runtimeCallbackToken(runtime *WorkspaceRuntime) string { + if runtime == nil { + return "" } - return "" + s.workspaceMu.RLock() + defer s.workspaceMu.RUnlock() + return strings.TrimSpace(runtime.CallbackToken) } func (s *Server) applyDetectedContainerUser(runtime *WorkspaceRuntime, detected string) { @@ -83,7 +105,7 @@ func (s *Server) provisionWorkspaceRuntime(ctx context.Context, runtime *Workspa return false, err } - callbackToken := strings.TrimSpace(runtime.CallbackToken) + callbackToken := s.runtimeCallbackToken(runtime) if callbackToken == "" { callbackToken = strings.TrimSpace(s.config.CallbackToken) } From 514647b68c383da0e4b605666bdaf3e7dc1b0be6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 09:58:40 +0000 Subject: [PATCH 12/19] refactor(vm-agent): move the publish job reporter out of mcp_build.go Move only, no behaviour change. mcp_build.go was already past the 500-line split threshold, and the token-source change grew it (rule 18). Co-Authored-By: Claude Opus 5.5 --- .../vm-agent/internal/server/mcp_build.go | 65 ---------------- .../internal/server/publish_job_reporter.go | 77 +++++++++++++++++++ 2 files changed, 77 insertions(+), 65 deletions(-) create mode 100644 packages/vm-agent/internal/server/publish_job_reporter.go diff --git a/packages/vm-agent/internal/server/mcp_build.go b/packages/vm-agent/internal/server/mcp_build.go index ff93524c12..b6501acd60 100644 --- a/packages/vm-agent/internal/server/mcp_build.go +++ b/packages/vm-agent/internal/server/mcp_build.go @@ -1,7 +1,6 @@ package server import ( - "bytes" "context" "encoding/json" "errors" @@ -364,70 +363,6 @@ func (s *Server) runPreparedBuildAndPublish(ctx context.Context, prepared *prepa return result, nil } -// publishCallbackToken reads the workspace token for each publish callback, so a -// job that runs while the token is renewed uses the renewed token. The token -// captured when the job was accepted is the fallback if the runtime is gone. -func (s *Server) publishCallbackToken(prepared *preparedBuildPublish) func() string { - return func() string { - if token := s.workspaceCallbackToken(prepared.WorkspaceID); token != "" { - return token - } - return prepared.Token - } -} - -type publishJobReporter struct { - baseURL string - projectID string - jobID string - token func() string - client *http.Client - log *slog.Logger -} - -func newPublishJobReporter(baseURL, projectID, jobID string, token func() string, client *http.Client, log *slog.Logger) *publishJobReporter { - return &publishJobReporter{ - baseURL: strings.TrimRight(baseURL, "/"), - projectID: projectID, - jobID: jobID, - token: token, - client: client, - log: log.With("component", "publish-job-reporter", "publishJobId", jobID), - } -} - -func (r *publishJobReporter) Event(ctx context.Context, event publish.Event) { - if r == nil || r.client == nil { - return - } - if event.Level == "" { - event.Level = "info" - } - raw, err := json.Marshal(event) - if err != nil { - r.log.Warn("marshal publish job event failed", "error", err) - return - } - url := r.baseURL + "/api/projects/" + r.projectID + "/deployment-publish-jobs/" + r.jobID + "/events" - req, err := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(raw)) - if err != nil { - r.log.Warn("create publish job event request failed", "error", err) - return - } - req.Header.Set("Authorization", "Bearer "+r.token()) - req.Header.Set("Content-Type", "application/json") - resp, err := r.client.Do(req) - if err != nil { - r.log.Warn("send publish job event failed", "eventType", event.EventType, "error", err) - return - } - defer resp.Body.Close() - if resp.StatusCode >= 300 { - body, _ := io.ReadAll(io.LimitReader(resp.Body, 1024)) - r.log.Warn("publish job event rejected", "eventType", event.EventType, "status", resp.StatusCode, "body", string(body)) - } -} - // resolveBuildSourceDir returns the host filesystem path of the agent's actual // working tree so the build publishes the agent's committed source rather than // the boot-time host clone. diff --git a/packages/vm-agent/internal/server/publish_job_reporter.go b/packages/vm-agent/internal/server/publish_job_reporter.go new file mode 100644 index 0000000000..381ede0c31 --- /dev/null +++ b/packages/vm-agent/internal/server/publish_job_reporter.go @@ -0,0 +1,77 @@ +package server + +import ( + "bytes" + "context" + "encoding/json" + "io" + "log/slog" + "net/http" + "strings" + + "github.com/workspace/vm-agent/internal/publish" +) + +// publishCallbackToken reads the workspace token for each publish callback, so a +// job that runs while the token is renewed uses the renewed token. The token +// captured when the job was accepted is the fallback if the runtime is gone. +func (s *Server) publishCallbackToken(prepared *preparedBuildPublish) func() string { + return func() string { + if token := s.workspaceCallbackToken(prepared.WorkspaceID); token != "" { + return token + } + return prepared.Token + } +} + +type publishJobReporter struct { + baseURL string + projectID string + jobID string + token func() string + client *http.Client + log *slog.Logger +} + +func newPublishJobReporter(baseURL, projectID, jobID string, token func() string, client *http.Client, log *slog.Logger) *publishJobReporter { + return &publishJobReporter{ + baseURL: strings.TrimRight(baseURL, "/"), + projectID: projectID, + jobID: jobID, + token: token, + client: client, + log: log.With("component", "publish-job-reporter", "publishJobId", jobID), + } +} + +func (r *publishJobReporter) Event(ctx context.Context, event publish.Event) { + if r == nil || r.client == nil { + return + } + if event.Level == "" { + event.Level = "info" + } + raw, err := json.Marshal(event) + if err != nil { + r.log.Warn("marshal publish job event failed", "error", err) + return + } + url := r.baseURL + "/api/projects/" + r.projectID + "/deployment-publish-jobs/" + r.jobID + "/events" + req, err := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(raw)) + if err != nil { + r.log.Warn("create publish job event request failed", "error", err) + return + } + req.Header.Set("Authorization", "Bearer "+r.token()) + req.Header.Set("Content-Type", "application/json") + resp, err := r.client.Do(req) + if err != nil { + r.log.Warn("send publish job event failed", "eventType", event.EventType, "error", err) + return + } + defer resp.Body.Close() + if resp.StatusCode >= 300 { + body, _ := io.ReadAll(io.LimitReader(resp.Body, 1024)) + r.log.Warn("publish job event rejected", "eventType", event.EventType, "status", resp.StatusCode, "body", string(body)) + } +} From a04a7a155c35abb55bd0a96b9a53b6c24f924751 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 10:11:16 +0000 Subject: [PATCH 13/19] fix(api): share the hibernate token through a plain memo object The first version exported a factory from node-agent, and the session-sleep suites partially mock node-agent, so 29 of their tests failed with "No hibernateCallbackTokenDelivery export". The wait loop now passes a plain HibernateCallbackTokenDelivery object, imported as a type, so it adds no runtime import. The tests assert that one object serves every poll, that a shared object mints once, and that an already-minted token is delivered as is. All three are mutation-verified. Co-Authored-By: Claude Opus 5.5 --- .../services/node-agent-session-snapshots.ts | 27 ++++------ apps/api/src/services/node-agent.ts | 2 +- .../services/session-sleep-snapshot-wait.ts | 9 ++-- .../unit/session-sleep-snapshot-wait.test.ts | 13 +---- .../workspace-callback-token-renewal.test.ts | 49 +++++++++++++------ 5 files changed, 49 insertions(+), 51 deletions(-) diff --git a/apps/api/src/services/node-agent-session-snapshots.ts b/apps/api/src/services/node-agent-session-snapshots.ts index d32060b7ae..c20106a809 100644 --- a/apps/api/src/services/node-agent-session-snapshots.ts +++ b/apps/api/src/services/node-agent-session-snapshots.ts @@ -57,21 +57,14 @@ function requestSessionSnapshot( ); } -/** Supplies the workspace token one hibernate request delivers; null delivers none. */ -export type HibernateCallbackTokenDelivery = () => Promise; - /** - * Mint at most one delivered workspace token for repeats of the same hibernate request. - * The agent installs a delivered token as soon as it reads the request, accepted or not, - * so a caller that repeats the request until the agent accepts it resends the token it - * minted first instead of signing a new one on every poll. + * Lets repeats of one hibernate request, for one workspace and node, share the workspace + * token minted for the first. The agent installs a delivered token as soon as it reads the + * request, accepted or not, so a caller that repeats the request until the agent accepts + * it passes the same object every time instead of signing a new token on every poll. */ -export function hibernateCallbackTokenDelivery( - env: Env, - target: { workspaceId: string; nodeId: string } -): HibernateCallbackTokenDelivery { - let minted: Promise | undefined; - return () => (minted ??= mintWorkspaceCallbackTokenForNodeDelivery(env, target)); +export interface HibernateCallbackTokenDelivery { + minted?: Promise; } export async function hibernateAgentSessionOnNode( @@ -81,15 +74,13 @@ export async function hibernateAgentSessionOnNode( env: Env, userId: string, input: SessionSnapshotRequest, - deliverWorkspaceCallbackToken: HibernateCallbackTokenDelivery = hibernateCallbackTokenDelivery( - env, - { workspaceId, nodeId } - ) + delivery: HibernateCallbackTokenDelivery = {} ): Promise { // The capture's prepare/progress/complete/failure callbacks authenticate with the // workspace token the agent holds. Deliver a fresh one over this node-management // request, exactly as create/restore do, so a long-awake workspace can still sleep. - const workspaceCallbackToken = await deliverWorkspaceCallbackToken(); + delivery.minted ??= mintWorkspaceCallbackTokenForNodeDelivery(env, { workspaceId, nodeId }); + const workspaceCallbackToken = await delivery.minted; return requestSessionSnapshot('hibernate', nodeId, workspaceId, sessionId, env, userId, { ...input, ...(workspaceCallbackToken ? { workspaceCallbackToken } : {}), diff --git a/apps/api/src/services/node-agent.ts b/apps/api/src/services/node-agent.ts index 88a260b796..456882c670 100644 --- a/apps/api/src/services/node-agent.ts +++ b/apps/api/src/services/node-agent.ts @@ -755,9 +755,9 @@ export async function sendPromptToAgentOnNode( } } +export type { HibernateCallbackTokenDelivery } from './node-agent-session-snapshots'; export { hibernateAgentSessionOnNode, - hibernateCallbackTokenDelivery, restoreAgentSessionOnNode, } from './node-agent-session-snapshots'; diff --git a/apps/api/src/services/session-sleep-snapshot-wait.ts b/apps/api/src/services/session-sleep-snapshot-wait.ts index 58ba94135c..c9f8d8a4ba 100644 --- a/apps/api/src/services/session-sleep-snapshot-wait.ts +++ b/apps/api/src/services/session-sleep-snapshot-wait.ts @@ -4,7 +4,7 @@ import type * as schema from '../db/schema'; import type { Env } from '../env'; import { log } from '../lib/logger'; import { parsePositiveInt } from '../lib/route-helpers'; -import { hibernateAgentSessionOnNode, hibernateCallbackTokenDelivery } from './node-agent'; +import { hibernateAgentSessionOnNode,type HibernateCallbackTokenDelivery } from './node-agent'; import { completeActiveSessionSnapshotAsDegraded, getSessionSnapshotCaptureState, @@ -68,10 +68,7 @@ export async function waitForFinalSessionSnapshot( let lastProgressAt = Date.now(); let lastProgressToken = ''; // One delivered workspace token for every poll of this request, not one per poll. - const deliverWorkspaceCallbackToken = hibernateCallbackTokenDelivery(env, { - workspaceId: input.workspaceId, - nodeId: input.nodeId, - }); + const tokenDelivery: HibernateCallbackTokenDelivery = {}; while (Date.now() < requestDeadline) { const current = await getSessionSnapshotCaptureState(db, input.chatSessionId); @@ -96,7 +93,7 @@ export async function waitForFinalSessionSnapshot( agentType: input.agentType, background: true, }, - deliverWorkspaceCallbackToken + tokenDelivery )) as SnapshotResult & { accepted?: unknown }; if (result.status !== 'pending') { throw new Error(`Workspace snapshot request was not accepted (${String(result.status)})`); diff --git a/apps/api/tests/unit/session-sleep-snapshot-wait.test.ts b/apps/api/tests/unit/session-sleep-snapshot-wait.test.ts index 2cb7523165..878e420779 100644 --- a/apps/api/tests/unit/session-sleep-snapshot-wait.test.ts +++ b/apps/api/tests/unit/session-sleep-snapshot-wait.test.ts @@ -9,12 +9,10 @@ import { createSchemaTables, createSqliteD1 } from '../helpers/sqlite-d1'; const mocks = vi.hoisted(() => ({ hibernateAgentSessionOnNode: vi.fn(), - hibernateCallbackTokenDelivery: vi.fn(), })); vi.mock('../../src/services/node-agent', () => ({ hibernateAgentSessionOnNode: mocks.hibernateAgentSessionOnNode, - hibernateCallbackTokenDelivery: mocks.hibernateCallbackTokenDelivery, })); describe('waitForFinalSessionSnapshot', () => { @@ -109,9 +107,6 @@ describe('waitForFinalSessionSnapshot', () => { SESSION_SNAPSHOT_PROGRESS_IDLE_TIMEOUT_MS: '1000', SESSION_SNAPSHOT_POLL_INTERVAL_MS: '1', } as unknown as Env; - const delivery = vi.fn(async () => 'delivered-token'); - mocks.hibernateCallbackTokenDelivery.mockReset(); - mocks.hibernateCallbackTokenDelivery.mockReturnValueOnce(delivery); mocks.hibernateAgentSessionOnNode.mockReset(); mocks.hibernateAgentSessionOnNode .mockResolvedValueOnce({ status: 'pending', accepted: false }) @@ -127,15 +122,11 @@ describe('waitForFinalSessionSnapshot', () => { userId: 'user-1', }); - expect(mocks.hibernateCallbackTokenDelivery).toHaveBeenCalledTimes(1); - expect(mocks.hibernateCallbackTokenDelivery).toHaveBeenCalledWith(testEnv, { - workspaceId: 'workspace-1', - nodeId: 'node-1', - }); const calls = mocks.hibernateAgentSessionOnNode.mock.calls; expect(calls).toHaveLength(3); + expect(calls[0][6]).toEqual(expect.any(Object)); for (const call of calls) { - expect(call[6]).toBe(delivery); + expect(call[6]).toBe(calls[0][6]); } } finally { sqlite.close(); diff --git a/apps/api/tests/workers/workspace-callback-token-renewal.test.ts b/apps/api/tests/workers/workspace-callback-token-renewal.test.ts index 12155e14f4..bfdaab5691 100644 --- a/apps/api/tests/workers/workspace-callback-token-renewal.test.ts +++ b/apps/api/tests/workers/workspace-callback-token-renewal.test.ts @@ -24,7 +24,7 @@ import { } from '../../src/services/jwt'; import { hibernateAgentSessionOnNode, - hibernateCallbackTokenDelivery, + type HibernateCallbackTokenDelivery, } from '../../src/services/node-agent-session-snapshots'; import { mintWorkspaceCallbackTokenForNodeDelivery } from '../../src/services/workspace-callback-token-binding'; import { @@ -519,24 +519,44 @@ describe('hibernateAgentSessionOnNode workspace token delivery', () => { }); it('mints once for every repeat of a request that shares a delivery', async () => { - const deliver = hibernateCallbackTokenDelivery(testEnv, { - workspaceId: WS_ACTIVE, - nodeId: NODE_ID, - }); - - const first = deliver(); - expect(deliver()).toBe(first); + fetchMock.mockImplementation( + async () => + new Response(JSON.stringify({ status: 'pending', accepted: false }), { status: 202 }) + ); + vi.stubGlobal('fetch', fetchMock); + const delivery: HibernateCallbackTokenDelivery = {}; + const poll = () => + hibernateAgentSessionOnNode( + NODE_ID, + WS_ACTIVE, + 'agent-session-1', + testEnv, + USER_ID, + { chatSessionId: 'chat-1', runtime: 'vm', background: true }, + delivery + ); + + await poll(); + const minted = delivery.minted; + expect(minted).toBeInstanceOf(Promise); + await poll(); + + expect(delivery.minted).toBe(minted); + const tokens = fetchMock.mock.calls.map( + ([, init]) => JSON.parse(String((init as RequestInit).body)).workspaceCallbackToken + ); + expect(tokens).toEqual([await minted, await minted]); await expect( - verifyCallbackToken((await first) as string, testEnv, { expectedScope: 'workspace' }) + verifyCallbackToken(tokens[0] as string, testEnv, { expectedScope: 'workspace' }) ).resolves.toMatchObject({ workspace: WS_ACTIVE }); }); - it('delivers the token its caller supplies', async () => { - fetchMock.mockResolvedValue( - new Response(JSON.stringify({ status: 'pending', accepted: false }), { status: 202 }) + it('delivers the token already minted for an earlier poll', async () => { + fetchMock.mockImplementation( + async () => + new Response(JSON.stringify({ status: 'pending', accepted: false }), { status: 202 }) ); vi.stubGlobal('fetch', fetchMock); - const deliver = vi.fn(async () => 'token-minted-for-the-first-poll'); await hibernateAgentSessionOnNode( NODE_ID, @@ -545,10 +565,9 @@ describe('hibernateAgentSessionOnNode workspace token delivery', () => { testEnv, USER_ID, { chatSessionId: 'chat-1', runtime: 'vm', background: true }, - deliver + { minted: Promise.resolve('token-minted-for-the-first-poll') } ); - expect(deliver).toHaveBeenCalledTimes(1); const [, init] = fetchMock.mock.calls[0] as [string, RequestInit]; expect(JSON.parse(String(init.body)).workspaceCallbackToken).toBe( 'token-minted-for-the-first-poll' From da5c1688a90e9fa76696515d613f743e56becb46 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 10:11:16 +0000 Subject: [PATCH 14/19] task: record review outcomes, evidence and follow-ups for token renewal Co-Authored-By: Claude Opus 5.5 --- ...-10-04-workspace-callback-token-renewal.md | 117 ++++++++++++++++-- ...generation-aware-callback-token-renewal.md | 26 ++++ ...10-04-snapshot-relay-node-proof-in-body.md | 24 ++++ ...-after-bootstrap-workspace-token-writer.md | 30 +++++ 4 files changed, 189 insertions(+), 8 deletions(-) create mode 100644 tasks/backlog/2026-10-04-instant-generation-aware-callback-token-renewal.md create mode 100644 tasks/backlog/2026-10-04-snapshot-relay-node-proof-in-body.md create mode 100644 tasks/backlog/2026-10-04-update-after-bootstrap-workspace-token-writer.md diff --git a/tasks/active/2026-10-04-workspace-callback-token-renewal.md b/tasks/active/2026-10-04-workspace-callback-token-renewal.md index 911e4938f0..e8b2b3986a 100644 --- a/tasks/active/2026-10-04-workspace-callback-token-renewal.md +++ b/tasks/active/2026-10-04-workspace-callback-token-renewal.md @@ -103,9 +103,39 @@ gets `401 Invalid or expired callback token` on every workspace-scoped callback. 5. **ACP SessionHost:** read the callback token through a lock-free accessor that renewal updates (rule 46: nothing reachable from the ACP notification goroutine may take `mu`). +Changes from local review (security, Go, Cloudflare, test, task-completion, docs, env): + +6. **VM-only renewal.** The route refuses Instant (cf-container) workspaces, like delivery. A + container generation is replaced under the same nodeId and the route cannot tell a superseded + generation from the current one, so renewing would let a surviving old generation extend its + workspace authority. Instant containers get a fresh token per cold wake, as before. Follow-up: + `tasks/backlog/2026-10-04-instant-generation-aware-callback-token-renewal.md`. +7. **Rate limit (rule 28 §4).** Authenticated renewal attempts count against + `RATE_LIMIT_CALLBACK_TOKEN_RENEWAL` (default 12 per `RATE_LIMIT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS` + = 3600) per workspace, in one guarded D1 upsert (migration 0179, cascade-deleted with the + workspace). A slot is spent only after both proofs, the node binding and the active check, so a + caller without the credentials cannot use up the legitimate agent's quota. 429 + `Retry-After`; + the agent backs off. +8. **Every consumer reads the current token.** A new concurrency test found that + `callbackTokenForWorkspace`/`workspaceCallbackToken` read `runtime.CallbackToken` through a shared + pointer outside `workspaceMu` while renewal writes it under the lock (a real data race). All + readers now take the lock (SessionHost creation, git credential auth, publish, provisioning, + standalone clone). Publish jobs (up to `DeployBuildPublishTimeout`) read the current token per + request instead of the one captured at start. +9. **Identity check on install.** The agent never installs a renewed or delivered token whose claims + name another workspace or a node (claims read unverified; the control plane verifies). A renewal + response like that is retried after backoff, not latched. +10. **One delivered token per hibernate wait.** The sleep wait loop repeats the hibernate request + every poll until the agent accepts it, and the agent installs a delivered token whether or not + it accepts. One token minted for the first poll now serves every poll (binding checked at that + mint, before the first delivery), instead of up to ~300 signs and 600 D1 reads. +11. The agent's ratio clamp matches the control plane: non-finite means the default, finite values + clamp to 0.1-0.9. + ## Implementation checklist ### API + - [x] `jwt.ts`: renewal signing preserves generation (`gen_iat`); claim readers live in `callback-token-claims.ts` (tests partially mock `jwt.ts`); stale guard reads `gen_iat` first - [x] New service `workspace-callback-token-renewal.ts`: dual-proof verification, D1 binding checks @@ -119,8 +149,14 @@ gets `401 Invalid or expired callback token` on every workspace-scoped callback. - [x] `node-agent-session-snapshots.ts`: include fresh token on hibernate when bound + active (VM only) - [x] Env vars documented: env-reference skill (API `CALLBACK_TOKEN_*`, agent `WORKSPACE_CALLBACK_TOKEN_*`, `MSG_AUTH_RENEWAL_WAIT`), public VM agent reference +- [x] Review: atomic per-workspace renewal limit (migration 0179, `workspace-callback-token-renewal-rate-limit.ts`, + `RATE_LIMIT_CALLBACK_TOKEN_RENEWAL[_WINDOW_SECONDS]`, 429 + Retry-After) +- [x] Review: renewal refuses Instant runtimes (`isInstantRuntimeBinding`, shared with delivery) +- [x] Review: one delivered token per hibernate wait (`HibernateCallbackTokenDelivery`, a plain memo object + so the partial `node-agent` mocks in the sleep suites keep working) ### VM agent + - [x] Config: renewal ratio (clamped), retry initial/max, request timeout - [x] `workspace_callback_token_renewal.go`: due selection with injected clock, request with both tokens, response classification, bounded backoff, rejection latch keyed on the token, CAS apply @@ -129,8 +165,14 @@ gets `401 Invalid or expired callback token` on every workspace-scoped callback. - [x] `acp.SessionHost`: lock-free current-token accessor + `SetCallbackToken`; replace reads - [x] `messagereport`: stale-token retry, park-on-401 without deleting rows, resume on new token, pause beyond `MSG_AUTH_RENEWAL_WAIT` reported once (slog.Error + node errorreport) +- [x] Review: all `runtime.CallbackToken` reads under `workspaceMu` (data race found by the new + host-creation test) +- [x] Review: publish job reporter and publish control plane read the current token per request +- [x] Review: refuse renewed/delivered tokens naming another workspace or a node +- [x] Review: ratio clamp aligned with the API; `mcp_build.go` split (move-only) under rule 18 ### Tests + - [x] Workers test through the real route (29): renew success (gen preserved, scope, new exp), expired denied, node-as-workspace and workspace-as-node proofs denied, foreign/moved node, owner mismatch, deleted/stopped/terminal node 410, missing row 410, claim/path mismatch, not-due, @@ -143,8 +185,17 @@ gets `401 Invalid or expired callback token` on every workspace-scoped callback. propagation to reporter + every host of the workspace only; restart hydration; reporter park/resume/stale-401/no-duplicate/budget/bounded; hibernate handler uses delivered token - [x] `go test -race` for server, messagereport, acp, config +- [x] Review additions. API: rate limit at-limit/rollover/per-workspace/concurrent/missing-workspace + on real SQLite; route-level 429 with Retry-After, where failed-auth attempts do not spend quota; + cascade delete in Miniflare D1; Instant refusal (unit + route); move race during renewal; quota + race; one delivery per hibernate wait (unit), delivery memoization and supplied token (workers). + Agent: 429 backs off; foreign/node token never installed (renewal + delivery); reporter → + node error reporter wiring through `getOrCreateReporter`; no token in logs across every outcome; + SessionHost created during a token change ends with the new token (200 rounds, -race); publish + job events and publish control plane use a token renewed mid-job ### Docs / rollout + - [x] Public docs: security architecture (callback token scopes, lifetime, renewal), VM agent env - [x] Rollout note (PR body): API push heals old agents for snapshot calls only (proven against pinned 7a9782c90 source); other consumers need new nodes; no hot replacement; AI proxy env token @@ -152,14 +203,22 @@ gets `401 Invalid or expired callback token` on every workspace-scoped callback. ## Acceptance criteria -- [ ] A VM workspace awake past 24h keeps working: snapshot prepare/progress/complete, git-token, - runtime assets, messages, ACP activity/usage/interactions (new agent) -- [ ] Old agents: hibernate/snapshot callbacks succeed after 24h via the API push alone -- [ ] Renewal never mints for: expired tokens, node-only callers, foreign nodes, moved/deleted/ - stopped workspaces, terminal nodes, superseded Instant generations -- [ ] No message loss or duplication across a token rotation; 401 handling is bounded -- [ ] No tokens in logs; responses carrying tokens are `no-store` -- [ ] Independent security review passes; parent informed of the design +- [x] A VM workspace awake past 24h keeps working: snapshot prepare/progress/complete, git-token, + runtime assets, messages, ACP activity/usage/interactions, publish jobs (new agent). + Evidence: `TestWorkspaceTokenRenewal_RenewsAtTheRefreshPointAndKeepsAWorkspaceAlivePast24h`, + `..._ReachesParkedReporterAndEverySessionHostOfTheWorkspace`, `TestPublishJob_CallbacksUseATokenRenewedWhileTheJobRuns`, + workers "lets a workspace callback that failed after 24h succeed with the renewed token" +- [x] Old agents: hibernate/snapshot callbacks succeed after 24h via the API push alone. + Evidence: `workspace_callback_token_hibernate_test.go` passes unchanged on pinned 7a9782c90 + (re-verified independently by the test review); workers hibernate-body tests +- [x] Renewal never mints for: expired tokens, node-only callers, foreign nodes, moved/deleted/ + stopped workspaces, terminal nodes, any Instant generation (Instant renewal refused entirely), + or past the per-workspace rate limit. Evidence: workers suite (34) + unit suite (17) +- [x] No message loss or duplication across a token rotation; 401 handling is bounded. + Evidence: `credential_test.go` (7), server dedupe by messageId, outbox cap +- [x] No tokens in logs; responses carrying tokens are `no-store`. + Evidence: `TestWorkspaceTokenRenewal_NeverLogsATokenValue`, error report body check, route test +- [ ] Independent security review passes; parent informed of the design (delta re-review pending) ## Staging @@ -178,6 +237,48 @@ Agent A1 expiry ordering, A2 CAS, A3 refusal latch, A4/A5 propagation, A6 persis credential retry; reporter R1/R2/R4/R5/R6; SessionHost S1; upsert publish U1. R3 (resume bookkeeping) is not a guard: parking is keyed on the rejected token, so a new token resumes by construction. +Review additions, each reverted once with the named test going red: + +- M11 Instant refusal removed → unit "refuses an Instant (cf-container) workspace, while the VM + control renews" and workers "refuses an Instant (cf-container) workspace..." +- M12 limit not enforced → workers "limits renewal attempts per workspace, counting only + authenticated ones" +- M13 quota consumed before authentication → same workers test (failed-auth attempts used up the + quota and the first valid renewal got 429) +- M14 delivery not memoized (`??=` → `=`) → workers "mints once for every repeat of a request + that shares a delivery" and "delivers the token already minted for an earlier poll" +- W1 a new delivery object per poll → unit "gives every poll of one hibernate request the same + workspace token delivery" +- G1 adopt identity check removed → `TestWorkspaceTokenDelivery_IgnoresTokensThatAreNotThisWorkspaces` +- G2 renewal identity check removed → `TestWorkspaceTokenRenewal_IgnoresARenewedTokenForAnotherWorkspace` +- G3 publish reporter uses captured token → `TestPublishJob_CallbacksUseATokenRenewedWhileTheJobRuns` +- G4 publish control plane ignores the token source → `TestHTTPControlPlaneUsesTheCurrentTokenForEachRequest` +- G5 token read outside `workspaceMu` → `TestWorkspaceTokenRenewal_HostCreatedDuringATokenChangeGetsTheNewToken` + (race detector) +- The quota's `workspace_missing` branch is covered at the limiter level only: at the service level + the post-sign identity re-read also returns 410, so the branch saves a signature rather than + changing the outcome. + +## Review results + +| Reviewer | Outcome | +| ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| security-auditor (full diff) | PASS-WITH-FINDINGS, no new trust boundary. HIGH rate limit fixed; MEDIUM Instant renewal fixed (VM-only); LOW move-race test added; LOW relay header filed as backlog; LOW cf-container token co-location tracked by Idea 01M432G3276YZWCP3HEJ5B25J5 | +| go-specialist | No CRITICAL/HIGH. MEDIUM identity check added; MEDIUM host-creation race test added (it found the token-read data race, fixed); MEDIUM O(N) host scan per token change left as a residual (one pass per renewal, ~12h apart); LOW test clock hardened | +| cloudflare-specialist | Approve. MEDIUM per-poll re-mint fixed; LOW doc pointer fixed; LOW rate limit fixed | +| test-engineer | MEDIUM Instant route test (now a refusal test); MEDIUM reporter wiring test added; LOW `UpdateAfterBootstrap` filed as backlog | +| task-completion-validator | HIGH publish-job custody fixed (live token source); LOW no-token-in-logs test added; ACs checked | +| doc-sync-validator | Rule 54 updated; `security.md` and api-reference wording fixed; AC3 conflict resolved by VM-only renewal | +| env-validator | MEDIUM clamp aligned; LOW `.env.example` entries added | +| constitution-validator | PASS | + +## Follow-ups + +- `tasks/backlog/2026-10-04-snapshot-relay-node-proof-in-body.md` +- `tasks/backlog/2026-10-04-update-after-bootstrap-workspace-token-writer.md` +- `tasks/backlog/2026-10-04-instant-generation-aware-callback-token-renewal.md` +- SAM Idea 01M432G3276YZWCP3HEJ5B25J5 (credentials baked into a running agent process) + ## References - `.claude/rules/28-credential-resolution-fallback-tests.md`, `.claude/rules/34` (callback auth), diff --git a/tasks/backlog/2026-10-04-instant-generation-aware-callback-token-renewal.md b/tasks/backlog/2026-10-04-instant-generation-aware-callback-token-renewal.md new file mode 100644 index 0000000000..8c197dd24f --- /dev/null +++ b/tasks/backlog/2026-10-04-instant-generation-aware-callback-token-renewal.md @@ -0,0 +1,26 @@ +# Renew Instant workspace callback tokens through the container DO + +## Problem + +Instant (cf-container) workspaces do not renew their workspace callback token. The renewal route +refuses them on purpose: a container generation is replaced under the same nodeId, and the route +cannot tell a superseded generation from the current one. A container gets a fresh token on every +cold wake, and containers sleep after `CF_CONTAINER_SLEEP_AFTER` (default 1h) idle. So only an +Instant session kept awake for more than `CALLBACK_TOKEN_EXPIRY_MS` (24h) without sleeping still +hits expired-token 401s, exactly as before the renewal work. + +## Context + +Decided in `tasks/archive/2026-10-04-workspace-callback-token-renewal.md` (security review MEDIUM). +The `VmAgentContainer` DO always knows its current generation and already talks to it over a +trusted channel (`containerFetch` with a node-management token). That channel can push a fresh +token to the current container only, which makes renewal generation-aware by construction. + +## Acceptance Criteria + +- [ ] Measure first: how often Instant sessions stay awake past 24h in production (Workers Logs + 401s from cf-container nodes); close this task if it never happens +- [ ] If needed, the DO pushes a fresh workspace token (keeping `gen_iat`) to its current container + before expiry, on its existing keepalive alarm +- [ ] A superseded generation never receives a token, proven by a test that replaces the + generation and checks that the old one gets nothing diff --git a/tasks/backlog/2026-10-04-snapshot-relay-node-proof-in-body.md b/tasks/backlog/2026-10-04-snapshot-relay-node-proof-in-body.md new file mode 100644 index 0000000000..eae404e855 --- /dev/null +++ b/tasks/backlog/2026-10-04-snapshot-relay-node-proof-in-body.md @@ -0,0 +1,24 @@ +# Move the snapshot upload relay's node proof out of a custom header + +## Problem + +`verifySessionSnapshotRelayAuthorization` (`apps/api/src/services/session-snapshot-upload-relay.ts`) +takes its second credential, a node-scoped callback token, in the custom header +`X-SAM-Relay-Authorization` (`SESSION_SNAPSHOT_RELAY_AUTHORIZATION_HEADER`). Workers Logs records +request headers. Cloudflare documents masking for `Authorization`, but not for custom headers, so +this bearer token may be stored in plain text in logs. + +## Context + +Found by the security review of workspace callback-token renewal +(`tasks/archive/2026-10-04-workspace-callback-token-renewal.md`). The renewal route sends the same +kind of node proof in the JSON body for exactly this reason. The relay route is older and was not +part of that change. + +## Acceptance Criteria + +- [ ] Check in Workers Logs (staging) whether `X-SAM-Relay-Authorization` values are recorded +- [ ] If recorded, or if that cannot be ruled out, carry the node proof in the request body and + keep accepting the header until every running VM agent sends the new shape (rule 54) +- [ ] Workers test through the real route for both shapes, plus a rejected-proof control +- [ ] Update the VM agent sender and the docs that describe the relay diff --git a/tasks/backlog/2026-10-04-update-after-bootstrap-workspace-token-writer.md b/tasks/backlog/2026-10-04-update-after-bootstrap-workspace-token-writer.md new file mode 100644 index 0000000000..35ac77b4b4 --- /dev/null +++ b/tasks/backlog/2026-10-04-update-after-bootstrap-workspace-token-writer.md @@ -0,0 +1,30 @@ +# Audit the workspace token write in UpdateAfterBootstrap + +## Problem + +`Server.UpdateAfterBootstrap` (`packages/vm-agent/internal/server/server.go`) writes +`cfg.CallbackToken` into the boot workspace's `WorkspaceRuntime.CallbackToken` directly. It +bypasses `adoptWorkspaceCallbackTokenLocked` (never adopt an earlier-expiring token, never adopt +another workspace's or a node-scoped token) and the persist-then-propagate step that every other +writer of the runtime token now uses. It also refreshes message reporter tokens +(`setTokenAllReporters`) before the write, so reporters get the old value. + +It is unclear which token `cfg.CallbackToken` holds at that point. The bootstrap token data carries +a callback token, and if that is node-scoped, the boot workspace would hold a node token that every +workspace callback rejects. + +## Context + +Found by the test review of workspace callback-token renewal +(`tasks/archive/2026-10-04-workspace-callback-token-renewal.md`). That task enumerated four writers +of the runtime token and routed the other three through the adopt/persist/propagate path. This one +runs once at boot, before any token is old enough to renew, so it does not race renewal in +practice. + +## Acceptance Criteria + +- [ ] Establish which token scope `cfg.CallbackToken` holds when `UpdateAfterBootstrap` runs, and + for which node types `cfg.WorkspaceID` is set +- [ ] Route the write through the shared adopt/persist/propagate path, or remove it if it is dead +- [ ] Test that drives `UpdateAfterBootstrap` and asserts the runtime, the reporter and every + SessionHost end up with the same token From 22e06a4cf3289e5162c24356d7dd9537b257eb10 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 10:25:53 +0000 Subject: [PATCH 15/19] refactor(vm-agent): move PTY and container-resolver helpers out of workspace_routing.go Move only, byte-identical function bodies. workspace_routing.go had reached 798 lines, and this branch adds to it (rule 18: split when adding to a file over 500 lines). It is now 603 lines; the helpers live in workspace_pty.go. Co-Authored-By: Claude Opus 5.5 --- .../vm-agent/internal/server/workspace_pty.go | 203 ++++++++++++++++++ .../internal/server/workspace_routing.go | 194 ----------------- 2 files changed, 203 insertions(+), 194 deletions(-) create mode 100644 packages/vm-agent/internal/server/workspace_pty.go diff --git a/packages/vm-agent/internal/server/workspace_pty.go b/packages/vm-agent/internal/server/workspace_pty.go new file mode 100644 index 0000000000..a635b029d8 --- /dev/null +++ b/packages/vm-agent/internal/server/workspace_pty.go @@ -0,0 +1,203 @@ +package server + +// PTY manager and devcontainer resolver construction for workspace runtimes. + +import ( + "context" + "fmt" + "strings" + + "github.com/workspace/vm-agent/internal/container" + "github.com/workspace/vm-agent/internal/pty" +) + +func (s *Server) newPTYManagerForWorkspace( + workspaceID, + workspaceDir, + containerWorkDir, + containerLabelValue, + containerUser string, +) *pty.Manager { + workDir := workspaceDir + if s.config.ContainerMode { + workDir = containerWorkDir + } + resolvedContainerUser := strings.TrimSpace(containerUser) + if resolvedContainerUser == "" { + resolvedContainerUser = strings.TrimSpace(s.config.ContainerUser) + } + + config := pty.ManagerConfig{ + DefaultShell: s.config.DefaultShell, + DefaultRows: s.config.DefaultRows, + DefaultCols: s.config.DefaultCols, + WorkDir: workDir, + ContainerResolver: s.ptyManagerContainerResolverForLabel(containerLabelValue), + ContainerUser: resolvedContainerUser, + GracePeriod: s.config.PTYOrphanGracePeriod, + BufferSize: s.config.PTYOutputBufferSize, + SessionIDMaxLength: s.config.TerminalSessionIDMaxLength, + CloseGrace: s.config.PTYCloseGracePeriod, + } + + manager := pty.NewManager(config) + if s.shouldReusePrimaryPTYManager(workspaceID, workspaceDir, containerWorkDir, containerLabelValue) { + return s.ptyManager + } + + return manager +} + +func (s *Server) shouldReusePrimaryPTYManager(workspaceID, workspaceDir, containerWorkDir, containerLabelValue string) bool { + if s == nil || s.ptyManager == nil { + return false + } + + // Preserve compatibility with legacy single-workspace host mode. + if !s.config.ContainerMode && len(s.workspaces) == 0 { + return true + } + + configuredWorkspaceID := strings.TrimSpace(s.config.WorkspaceID) + if configuredWorkspaceID == "" || strings.TrimSpace(workspaceID) != configuredWorkspaceID { + return false + } + + expectedWorkspaceDir := strings.TrimSpace(s.workspaceDirForRuntime(configuredWorkspaceID)) + if expectedWorkspaceDir == "" { + expectedWorkspaceDir = "/workspace" + } + if strings.TrimSpace(workspaceDir) != expectedWorkspaceDir { + return false + } + + if !s.config.ContainerMode { + return true + } + + expectedContainerLabel := strings.TrimSpace(s.config.ContainerLabelValue) + if expectedContainerLabel == "" { + expectedContainerLabel = expectedWorkspaceDir + } + if strings.TrimSpace(containerLabelValue) != expectedContainerLabel { + return false + } + + expectedContainerWorkDir := strings.TrimSpace(s.config.ContainerWorkDir) + if expectedContainerWorkDir == "" { + expectedContainerWorkDir = deriveContainerWorkDirForRepo(expectedWorkspaceDir, s.config.Repository) + } + if strings.TrimSpace(containerWorkDir) != expectedContainerWorkDir { + return false + } + + return true +} + +func (s *Server) rebuildWorkspacePTYManager(runtime *WorkspaceRuntime) { + if runtime == nil { + return + } + if runtime.PTY != nil && runtime.PTY.SessionCount() > 0 { + return + } + runtime.PTY = s.newPTYManagerForWorkspace( + runtime.ID, + strings.TrimSpace(runtime.WorkspaceDir), + strings.TrimSpace(runtime.ContainerWorkDir), + strings.TrimSpace(runtime.ContainerLabelValue), + strings.TrimSpace(runtime.ContainerUser), + ) +} + +// pty.Manager does not expose its resolver, so we derive from config. +func (s *Server) ptyManagerContainerResolver() pty.ContainerResolver { + if !s.config.ContainerMode { + return nil + } + return s.ptyManagerContainerResolverFromConfig() +} + +func (s *Server) ptyManagerContainerResolverFromConfig() pty.ContainerResolver { + return s.ptyManagerContainerResolverForLabel(s.config.ContainerLabelValue) +} + +func (s *Server) ptyManagerContainerResolverForLabel(labelValue string) pty.ContainerResolver { + resolver := s.ptyManagerContainerResolverForLabelContext(labelValue) + if resolver == nil { + return nil + } + return func() (string, error) { + return resolver(context.Background()) + } +} + +func (s *Server) ptyManagerContainerResolverForLabelContext(labelValue string) func(context.Context) (string, error) { + if !s.config.ContainerMode { + return nil + } + + requestedLabel := strings.TrimSpace(labelValue) + labelCandidates := []string{} + if requestedLabel != "" { + // Workspace-scoped lookups must be strict to avoid cross-workspace routing + // when multiple containers share repo-derived or legacy label values. + labelCandidates = containerLabelCandidates(requestedLabel) + } else { + labelCandidates = containerLabelCandidates( + s.config.ContainerLabelValue, + s.config.WorkspaceDir, + "/workspace", + ) + } + if len(labelCandidates) == 0 { + return nil + } + + discoveries := make([]*container.Discovery, 0, len(labelCandidates)) + for _, candidate := range labelCandidates { + discoveries = append(discoveries, container.NewDiscovery(container.Config{ + LabelKey: s.config.ContainerLabelKey, + LabelValue: candidate, + CacheTTL: s.config.ContainerCacheTTL, + })) + } + + return func(ctx context.Context) (string, error) { + if ctx == nil { + ctx = context.Background() + } + var lastErr error + for _, discovery := range discoveries { + containerID, err := discovery.GetContainerIDContext(ctx) + if err == nil { + return containerID, nil + } + if ctxErr := ctx.Err(); ctxErr != nil { + return "", ctxErr + } + lastErr = err + } + if lastErr != nil { + return "", lastErr + } + return "", fmt.Errorf("no container label candidates configured") + } +} + +func containerLabelCandidates(values ...string) []string { + candidates := make([]string, 0, len(values)) + seen := make(map[string]struct{}, len(values)) + for _, value := range values { + trimmed := strings.TrimSpace(value) + if trimmed == "" { + continue + } + if _, ok := seen[trimmed]; ok { + continue + } + seen[trimmed] = struct{}{} + candidates = append(candidates, trimmed) + } + return candidates +} diff --git a/packages/vm-agent/internal/server/workspace_routing.go b/packages/vm-agent/internal/server/workspace_routing.go index 147b2e5776..0c8848d735 100644 --- a/packages/vm-agent/internal/server/workspace_routing.go +++ b/packages/vm-agent/internal/server/workspace_routing.go @@ -4,7 +4,6 @@ import ( "context" "crypto/rand" "encoding/hex" - "fmt" "log/slog" "net/http" "path/filepath" @@ -12,9 +11,7 @@ import ( "time" "github.com/workspace/vm-agent/internal/agentsessions" - "github.com/workspace/vm-agent/internal/container" "github.com/workspace/vm-agent/internal/eventstore" - "github.com/workspace/vm-agent/internal/pty" ) // firstNonEmpty returns the first non-empty string argument, or "". @@ -435,105 +432,6 @@ func (s *Server) upsertWorkspaceRuntime(workspaceID, repository, branch, status, return runtime } -func (s *Server) newPTYManagerForWorkspace( - workspaceID, - workspaceDir, - containerWorkDir, - containerLabelValue, - containerUser string, -) *pty.Manager { - workDir := workspaceDir - if s.config.ContainerMode { - workDir = containerWorkDir - } - resolvedContainerUser := strings.TrimSpace(containerUser) - if resolvedContainerUser == "" { - resolvedContainerUser = strings.TrimSpace(s.config.ContainerUser) - } - - config := pty.ManagerConfig{ - DefaultShell: s.config.DefaultShell, - DefaultRows: s.config.DefaultRows, - DefaultCols: s.config.DefaultCols, - WorkDir: workDir, - ContainerResolver: s.ptyManagerContainerResolverForLabel(containerLabelValue), - ContainerUser: resolvedContainerUser, - GracePeriod: s.config.PTYOrphanGracePeriod, - BufferSize: s.config.PTYOutputBufferSize, - SessionIDMaxLength: s.config.TerminalSessionIDMaxLength, - CloseGrace: s.config.PTYCloseGracePeriod, - } - - manager := pty.NewManager(config) - if s.shouldReusePrimaryPTYManager(workspaceID, workspaceDir, containerWorkDir, containerLabelValue) { - return s.ptyManager - } - - return manager -} - -func (s *Server) shouldReusePrimaryPTYManager(workspaceID, workspaceDir, containerWorkDir, containerLabelValue string) bool { - if s == nil || s.ptyManager == nil { - return false - } - - // Preserve compatibility with legacy single-workspace host mode. - if !s.config.ContainerMode && len(s.workspaces) == 0 { - return true - } - - configuredWorkspaceID := strings.TrimSpace(s.config.WorkspaceID) - if configuredWorkspaceID == "" || strings.TrimSpace(workspaceID) != configuredWorkspaceID { - return false - } - - expectedWorkspaceDir := strings.TrimSpace(s.workspaceDirForRuntime(configuredWorkspaceID)) - if expectedWorkspaceDir == "" { - expectedWorkspaceDir = "/workspace" - } - if strings.TrimSpace(workspaceDir) != expectedWorkspaceDir { - return false - } - - if !s.config.ContainerMode { - return true - } - - expectedContainerLabel := strings.TrimSpace(s.config.ContainerLabelValue) - if expectedContainerLabel == "" { - expectedContainerLabel = expectedWorkspaceDir - } - if strings.TrimSpace(containerLabelValue) != expectedContainerLabel { - return false - } - - expectedContainerWorkDir := strings.TrimSpace(s.config.ContainerWorkDir) - if expectedContainerWorkDir == "" { - expectedContainerWorkDir = deriveContainerWorkDirForRepo(expectedWorkspaceDir, s.config.Repository) - } - if strings.TrimSpace(containerWorkDir) != expectedContainerWorkDir { - return false - } - - return true -} - -func (s *Server) rebuildWorkspacePTYManager(runtime *WorkspaceRuntime) { - if runtime == nil { - return - } - if runtime.PTY != nil && runtime.PTY.SessionCount() > 0 { - return - } - runtime.PTY = s.newPTYManagerForWorkspace( - runtime.ID, - strings.TrimSpace(runtime.WorkspaceDir), - strings.TrimSpace(runtime.ContainerWorkDir), - strings.TrimSpace(runtime.ContainerLabelValue), - strings.TrimSpace(runtime.ContainerUser), - ) -} - // casWorkspaceStatus performs a compare-and-swap status transition. // It only sets the new status if the current status is one of the expected values. // Returns true if the transition was applied, false if the current status did not match. @@ -703,95 +601,3 @@ func randomEventID() string { _, _ = rand.Read(buf) return hex.EncodeToString(buf) } - -// pty.Manager does not expose its resolver, so we derive from config. -func (s *Server) ptyManagerContainerResolver() pty.ContainerResolver { - if !s.config.ContainerMode { - return nil - } - return s.ptyManagerContainerResolverFromConfig() -} - -func (s *Server) ptyManagerContainerResolverFromConfig() pty.ContainerResolver { - return s.ptyManagerContainerResolverForLabel(s.config.ContainerLabelValue) -} - -func (s *Server) ptyManagerContainerResolverForLabel(labelValue string) pty.ContainerResolver { - resolver := s.ptyManagerContainerResolverForLabelContext(labelValue) - if resolver == nil { - return nil - } - return func() (string, error) { - return resolver(context.Background()) - } -} - -func (s *Server) ptyManagerContainerResolverForLabelContext(labelValue string) func(context.Context) (string, error) { - if !s.config.ContainerMode { - return nil - } - - requestedLabel := strings.TrimSpace(labelValue) - labelCandidates := []string{} - if requestedLabel != "" { - // Workspace-scoped lookups must be strict to avoid cross-workspace routing - // when multiple containers share repo-derived or legacy label values. - labelCandidates = containerLabelCandidates(requestedLabel) - } else { - labelCandidates = containerLabelCandidates( - s.config.ContainerLabelValue, - s.config.WorkspaceDir, - "/workspace", - ) - } - if len(labelCandidates) == 0 { - return nil - } - - discoveries := make([]*container.Discovery, 0, len(labelCandidates)) - for _, candidate := range labelCandidates { - discoveries = append(discoveries, container.NewDiscovery(container.Config{ - LabelKey: s.config.ContainerLabelKey, - LabelValue: candidate, - CacheTTL: s.config.ContainerCacheTTL, - })) - } - - return func(ctx context.Context) (string, error) { - if ctx == nil { - ctx = context.Background() - } - var lastErr error - for _, discovery := range discoveries { - containerID, err := discovery.GetContainerIDContext(ctx) - if err == nil { - return containerID, nil - } - if ctxErr := ctx.Err(); ctxErr != nil { - return "", ctxErr - } - lastErr = err - } - if lastErr != nil { - return "", lastErr - } - return "", fmt.Errorf("no container label candidates configured") - } -} - -func containerLabelCandidates(values ...string) []string { - candidates := make([]string, 0, len(values)) - seen := make(map[string]struct{}, len(values)) - for _, value := range values { - trimmed := strings.TrimSpace(value) - if trimmed == "" { - continue - } - if _, ok := seen[trimmed]; ok { - continue - } - seen[trimmed] = struct{}{} - candidates = append(candidates, trimmed) - } - return candidates -} From 74b81b0a19898bc743c2c6ed25fc89a2444e540d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 10:26:14 +0000 Subject: [PATCH 16/19] task: archive workspace callback token renewal Task-completion validator: PASS. Security review: full diff and delta PASS. Co-Authored-By: Claude Opus 5.5 --- ...-10-04-workspace-callback-token-renewal.md | 29 +++++++++++-------- 1 file changed, 17 insertions(+), 12 deletions(-) rename tasks/{active => archive}/2026-10-04-workspace-callback-token-renewal.md (86%) diff --git a/tasks/active/2026-10-04-workspace-callback-token-renewal.md b/tasks/archive/2026-10-04-workspace-callback-token-renewal.md similarity index 86% rename from tasks/active/2026-10-04-workspace-callback-token-renewal.md rename to tasks/archive/2026-10-04-workspace-callback-token-renewal.md index e8b2b3986a..0165648d42 100644 --- a/tasks/active/2026-10-04-workspace-callback-token-renewal.md +++ b/tasks/archive/2026-10-04-workspace-callback-token-renewal.md @@ -169,7 +169,8 @@ Changes from local review (security, Go, Cloudflare, test, task-completion, docs host-creation test) - [x] Review: publish job reporter and publish control plane read the current token per request - [x] Review: refuse renewed/delivered tokens naming another workspace or a node -- [x] Review: ratio clamp aligned with the API; `mcp_build.go` split (move-only) under rule 18 +- [x] Review: ratio clamp aligned with the API; `mcp_build.go` and `workspace_routing.go` split + (move-only, byte-identical) under rule 18 ### Tests @@ -218,7 +219,10 @@ Changes from local review (security, Go, Cloudflare, test, task-completion, docs Evidence: `credential_test.go` (7), server dedupe by messageId, outbox cap - [x] No tokens in logs; responses carrying tokens are `no-store`. Evidence: `TestWorkspaceTokenRenewal_NeverLogsATokenValue`, error report body check, route test -- [ ] Independent security review passes; parent informed of the design (delta re-review pending) +- [x] Independent security review passes; parent informed of the design. Evidence: full-diff + adversarial review (no new trust boundary; HIGH/MEDIUM fixed) and delta re-review of the fix + commits (PASS). The original parent 01M42WJSH7238RWH5ZH7TSSFZG was cancelled; its resumed + conversation (task 01M4357G5G3105A4ZEJJ1MK7SD) was informed (message 01M4369BEY4CQ8VHJC4B1A7K1F) ## Staging @@ -261,16 +265,17 @@ Review additions, each reverted once with the named test going red: ## Review results -| Reviewer | Outcome | -| ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| security-auditor (full diff) | PASS-WITH-FINDINGS, no new trust boundary. HIGH rate limit fixed; MEDIUM Instant renewal fixed (VM-only); LOW move-race test added; LOW relay header filed as backlog; LOW cf-container token co-location tracked by Idea 01M432G3276YZWCP3HEJ5B25J5 | -| go-specialist | No CRITICAL/HIGH. MEDIUM identity check added; MEDIUM host-creation race test added (it found the token-read data race, fixed); MEDIUM O(N) host scan per token change left as a residual (one pass per renewal, ~12h apart); LOW test clock hardened | -| cloudflare-specialist | Approve. MEDIUM per-poll re-mint fixed; LOW doc pointer fixed; LOW rate limit fixed | -| test-engineer | MEDIUM Instant route test (now a refusal test); MEDIUM reporter wiring test added; LOW `UpdateAfterBootstrap` filed as backlog | -| task-completion-validator | HIGH publish-job custody fixed (live token source); LOW no-token-in-logs test added; ACs checked | -| doc-sync-validator | Rule 54 updated; `security.md` and api-reference wording fixed; AC3 conflict resolved by VM-only renewal | -| env-validator | MEDIUM clamp aligned; LOW `.env.example` entries added | -| constitution-validator | PASS | +| Reviewer | Outcome | +| ----------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| security-auditor (full diff) | PASS-WITH-FINDINGS, no new trust boundary. HIGH rate limit fixed; MEDIUM Instant renewal fixed (VM-only); LOW move-race test added; LOW relay header filed as backlog; LOW cf-container token co-location tracked by Idea 01M432G3276YZWCP3HEJ5B25J5 | +| security-auditor (delta re-review of the fix commits) | PASS. HIGH and MEDIUM verified fixed; no new trust boundary. LOW: renewal backoff ignores `Retry-After` on 429. Deferred: agent backoff (1m→30m) bounds retries, and a healthy agent never reaches the limit | +| go-specialist | No CRITICAL/HIGH. MEDIUM identity check added; MEDIUM host-creation race test added (it found the token-read data race, fixed); MEDIUM O(N) host scan per token change left as a residual (one pass per renewal, ~12h apart); LOW test clock hardened | +| cloudflare-specialist | Approve. MEDIUM per-poll re-mint fixed; LOW doc pointer fixed; LOW rate limit fixed | +| test-engineer | MEDIUM Instant route test (now a refusal test); MEDIUM reporter wiring test added; LOW `UpdateAfterBootstrap` filed as backlog | +| task-completion-validator | First run WARN: HIGH publish-job custody, fixed with a live token source; LOW no-token-in-logs test added. Re-run after the fixes: PASS (all six checks, test counts reproduced) | +| doc-sync-validator | Rule 54 updated; `security.md` and api-reference wording fixed; AC3 conflict resolved by VM-only renewal | +| env-validator | MEDIUM clamp aligned; LOW `.env.example` entries added | +| constitution-validator | PASS | ## Follow-ups From 4297bfeb2b8152bdfcb0e346d6f35488d991eec9 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 10:51:53 +0000 Subject: [PATCH 17/19] fix(vm-agent): decode callback token claims instead of skipping verification SonarCloud flagged both jwt.ParseUnverified calls (go:S5659, CRITICAL). The agent never authenticates anyone with these claims. It reads its own token's iat/exp to schedule renewal, and the workspace/scope claims to refuse a token the control plane minted for another workspace. The control plane verifies every token it receives. A small payload decoder states that intent and calls no verification API, so nothing looks like a skipped signature check. Behaviour is unchanged: G1/G2 still go red when the identity check is removed. New test covers malformed tokens, padded base64 and fractional numeric dates. Also fixes two code smells (shadowed max, needless variable). Co-Authored-By: Claude Opus 5.5 --- .../acp/session_host_callback_token_test.go | 2 +- .../workspace_callback_token_renewal.go | 69 ++++++++++++------- .../workspace_callback_token_safety_test.go | 37 ++++++++++ 3 files changed, 82 insertions(+), 26 deletions(-) diff --git a/packages/vm-agent/internal/acp/session_host_callback_token_test.go b/packages/vm-agent/internal/acp/session_host_callback_token_test.go index 6853e09548..993fe58764 100644 --- a/packages/vm-agent/internal/acp/session_host_callback_token_test.go +++ b/packages/vm-agent/internal/acp/session_host_callback_token_test.go @@ -84,7 +84,7 @@ func TestSessionHostCallbackTokenIsRaceFreeAcrossGoroutines(t *testing.T) { case <-stop: return default: - if token := host.callbackToken(); token == "" { + if host.callbackToken() == "" { t.Error("reader observed an empty callback token") return } diff --git a/packages/vm-agent/internal/server/workspace_callback_token_renewal.go b/packages/vm-agent/internal/server/workspace_callback_token_renewal.go index bceb6177c2..f8e182319e 100644 --- a/packages/vm-agent/internal/server/workspace_callback_token_renewal.go +++ b/packages/vm-agent/internal/server/workspace_callback_token_renewal.go @@ -24,6 +24,7 @@ package server import ( "bytes" "context" + "encoding/base64" "encoding/json" "errors" "io" @@ -35,8 +36,6 @@ import ( "sync" "time" - "github.com/golang-jwt/jwt/v5" - "github.com/workspace/vm-agent/internal/messagereport" ) @@ -169,36 +168,56 @@ func callbackTokenRenewalDue(token string, now time.Time, ratio float64) bool { return !now.Before(issuedAt.Add(time.Duration(float64(lifetime) * ratio))) } +// callbackTokenClaims are the claims the agent reads from a workspace callback +// token it holds or is offered. +type callbackTokenClaims struct { + Workspace string `json:"workspace"` + Scope string `json:"scope"` + Subject string `json:"sub"` + IssuedAt float64 `json:"iat"` + ExpiresAt float64 `json:"exp"` +} + +// decodeCallbackTokenClaims decodes a callback token's payload and verifies +// nothing. Nothing in the agent authenticates anyone with these claims: they only +// schedule renewal and stop a token the control plane minted for another +// workspace from being installed. The control plane verifies every token it +// receives. A token that does not decode returns ok=false. +func decodeCallbackTokenClaims(token string) (claims callbackTokenClaims, ok bool) { + parts := strings.Split(token, ".") + if len(parts) != 3 { + return callbackTokenClaims{}, false + } + payload, err := base64.RawURLEncoding.DecodeString(strings.TrimRight(parts[1], "=")) + if err != nil { + return callbackTokenClaims{}, false + } + if err := json.Unmarshal(payload, &claims); err != nil { + return callbackTokenClaims{}, false + } + return claims, true +} + // callbackTokenNamesOtherWorkspace reports whether a token's own claims say it is // not workspaceID's workspace token: node-scoped, or bound to another workspace -// (internal/auth/jwt.go applies the same rule to tokens it verifies). Claims are -// read WITHOUT verifying the signature; this only keeps a control-plane mistake -// from being installed. A token whose claims cannot be read is left for the -// control plane to reject. +// (internal/auth/jwt.go applies the same rule to tokens it verifies). A token +// whose claims cannot be read is left for the control plane to reject. func callbackTokenNamesOtherWorkspace(token, workspaceID string) bool { - var claims struct { - jwt.RegisteredClaims - Workspace string `json:"workspace"` - Scope string `json:"scope"` - } - if _, _, err := jwt.NewParser().ParseUnverified(token, &claims); err != nil { + claims, ok := decodeCallbackTokenClaims(token) + if !ok { return false } return claims.Scope == "node" || claims.Workspace != workspaceID || (claims.Subject != "" && claims.Subject != workspaceID) } -// callbackTokenLifetime reads iat/exp WITHOUT verifying the signature. It only -// schedules renewal; the control plane verifies every token it receives. +// callbackTokenLifetime reads iat/exp to schedule renewal. func callbackTokenLifetime(token string) (issuedAt, expiresAt time.Time, ok bool) { - claims := jwt.RegisteredClaims{} - if _, _, err := jwt.NewParser().ParseUnverified(token, &claims); err != nil { - return time.Time{}, time.Time{}, false - } - if claims.IssuedAt == nil || claims.ExpiresAt == nil || !claims.ExpiresAt.After(claims.IssuedAt.Time) { + claims, ok := decodeCallbackTokenClaims(token) + if !ok || claims.IssuedAt <= 0 || claims.ExpiresAt <= claims.IssuedAt { return time.Time{}, time.Time{}, false } - return claims.IssuedAt.Time, claims.ExpiresAt.Time, true + return time.Unix(int64(claims.IssuedAt), 0), time.Unix(int64(claims.ExpiresAt), 0), true } func (s *Server) renewWorkspaceCallbackToken(candidate workspaceTokenRenewalCandidate, nodeToken string) { @@ -239,13 +258,13 @@ func (s *Server) renewWorkspaceCallbackToken(candidate workspaceTokenRenewalCand } } -func workspaceTokenRenewalBackoff(failures int, initial, max time.Duration) time.Duration { - delay := initial - for i := 1; i < failures && delay < max; i++ { +func workspaceTokenRenewalBackoff(failures int, initialDelay, maxDelay time.Duration) time.Duration { + delay := initialDelay + for i := 1; i < failures && delay < maxDelay; i++ { delay *= 2 } - if delay > max { - return max + if delay > maxDelay { + return maxDelay } return delay } diff --git a/packages/vm-agent/internal/server/workspace_callback_token_safety_test.go b/packages/vm-agent/internal/server/workspace_callback_token_safety_test.go index 0cf9b97196..9171989276 100644 --- a/packages/vm-agent/internal/server/workspace_callback_token_safety_test.go +++ b/packages/vm-agent/internal/server/workspace_callback_token_safety_test.go @@ -10,6 +10,7 @@ package server import ( "bytes" "context" + "encoding/base64" "io" "log/slog" "net/http" @@ -388,3 +389,39 @@ func TestPublishJob_CallbacksUseATokenRenewedWhileTheJobRuns(t *testing.T) { t.Fatalf("first publish callback used %q, want the token the job started with", eventAuth[0]) } } + +func TestDecodeCallbackTokenClaims(t *testing.T) { + encode := func(payload string) string { + return "eyJhbGciOiJIUzI1NiJ9." + base64.RawURLEncoding.EncodeToString([]byte(payload)) + ".c2ln" + } + issued := time.Date(2026, 10, 4, 8, 0, 0, 0, time.UTC) + + claims, ok := decodeCallbackTokenClaims(workspaceTestToken(t, renewalTestWorkspace, issued, 24*time.Hour)) + if !ok || claims.Workspace != renewalTestWorkspace || claims.Subject != renewalTestWorkspace || + claims.Scope != "workspace" || int64(claims.IssuedAt) != issued.Unix() || int64(claims.ExpiresAt) != issued.Add(24*time.Hour).Unix() { + t.Fatalf("decoded %+v ok=%v from a production-shaped token", claims, ok) + } + // JWT numeric dates may be fractional; padded base64 must decode too. + padded := "e30." + base64.URLEncoding.EncodeToString([]byte(`{"workspace":"ws-1","iat":1791108000.5,"exp":1791194400.5}`)) + ".c2ln" + if claims, ok := decodeCallbackTokenClaims(padded); !ok || claims.Workspace != "ws-1" || int64(claims.ExpiresAt) != 1791194400 { + t.Fatalf("padded/fractional token decoded as %+v ok=%v", claims, ok) + } + + for name, token := range map[string]string{ + "opaque string": "not-a-jwt", + "two segments": "a.b", + "payload not b64": "a.!!!.c", + "payload not json": encode("not json"), + "payload not claims": encode(`["array"]`), + } { + if _, ok := decodeCallbackTokenClaims(token); ok { + t.Fatalf("%s decoded", name) + } + if callbackTokenNamesOtherWorkspace(token, renewalTestWorkspace) { + t.Fatalf("%s was refused; an unreadable token is left for the control plane to reject", name) + } + if _, _, ok := callbackTokenLifetime(token); ok { + t.Fatalf("%s produced a lifetime", name) + } + } +} From ee6429018488366234b7fffe292980bc36266b63 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 11:29:56 +0000 Subject: [PATCH 18/19] fix(api): mark every renewal response no-store, including errors CodeRabbit nitpick: Cache-Control was set only on the success path. It is now set before anything can throw, so 400/401/403/410/429 responses from the global error handler carry it too. Asserted on the 401 and 429 paths; removing the early header turns both red. Co-Authored-By: Claude Opus 5.5 --- apps/api/src/routes/workspaces/callback-token-renewal.ts | 5 +++-- .../tests/workers/workspace-callback-token-renewal.test.ts | 2 ++ 2 files changed, 5 insertions(+), 2 deletions(-) diff --git a/apps/api/src/routes/workspaces/callback-token-renewal.ts b/apps/api/src/routes/workspaces/callback-token-renewal.ts index 4569952d02..e38b458e94 100644 --- a/apps/api/src/routes/workspaces/callback-token-renewal.ts +++ b/apps/api/src/routes/workspaces/callback-token-renewal.ts @@ -35,6 +35,9 @@ const RenewalRequestSchema = v.object({ const callbackTokenRenewalRoutes = new Hono<{ Bindings: Env }>(); callbackTokenRenewalRoutes.post('/:id/callback-token/renew', async (c) => { + // A success response carries a credential; no intermediary may store any response. + // Set first so error responses from the global handler carry it too. + c.header('Cache-Control', 'no-store'); const workspaceId = c.req.param('id'); const workspaceToken = extractBearerToken(c.req.header('Authorization')); @@ -65,8 +68,6 @@ callbackTokenRenewalRoutes.post('/:id/callback-token/renew', async (c) => { } throw err; } - // The response can carry a credential; no intermediary may store it. - c.header('Cache-Control', 'no-store'); return c.json(result); }); diff --git a/apps/api/tests/workers/workspace-callback-token-renewal.test.ts b/apps/api/tests/workers/workspace-callback-token-renewal.test.ts index bfdaab5691..c0a3dbd324 100644 --- a/apps/api/tests/workers/workspace-callback-token-renewal.test.ts +++ b/apps/api/tests/workers/workspace-callback-token-renewal.test.ts @@ -229,6 +229,7 @@ describe('POST /api/workspaces/:id/callback-token/renew', () => { const response = await renew(WS_ACTIVE, expired, { nodeId: NODE_ID, nodeToken }); expect(response.status).toBe(401); + expect(response.headers.get('Cache-Control')).toBe('no-store'); expect(await errorCode(response)).toBe('UNAUTHORIZED'); }); @@ -390,6 +391,7 @@ describe('POST /api/workspaces/:id/callback-token/renew', () => { const limited = await renew(WS_RATE, aged, { nodeId: NODE_ID, nodeToken }); expect(limited.status).toBe(429); + expect(limited.headers.get('Cache-Control')).toBe('no-store'); expect(await errorCode(limited)).toBe('RATE_LIMIT_EXCEEDED'); const retryAfter = Number(limited.headers.get('Retry-After')); expect(retryAfter).toBeGreaterThanOrEqual(1); From ef6e7968358a6e655617ff6f9c6732e6fdb3cf1d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20Titsworth-Morin?= Date: Sun, 4 Oct 2026 11:42:08 +0000 Subject: [PATCH 19/19] fix(api): renumber the renewal rate-limit migration to 0180 #2223 merged 0179_session_snapshot_sleep_episode.sql first, and the D1 migration ordering check rejects two migrations with one prefix. Also re-formats the wait-loop import the rebase left unformatted. Co-Authored-By: Claude Opus 5.5 --- ... => 0180_workspace_callback_token_renewal_rate_limits.sql} | 0 apps/api/src/services/session-sleep-snapshot-wait.ts | 2 +- tasks/archive/2026-10-04-workspace-callback-token-renewal.md | 4 ++-- 3 files changed, 3 insertions(+), 3 deletions(-) rename apps/api/src/db/migrations/{0179_workspace_callback_token_renewal_rate_limits.sql => 0180_workspace_callback_token_renewal_rate_limits.sql} (100%) diff --git a/apps/api/src/db/migrations/0179_workspace_callback_token_renewal_rate_limits.sql b/apps/api/src/db/migrations/0180_workspace_callback_token_renewal_rate_limits.sql similarity index 100% rename from apps/api/src/db/migrations/0179_workspace_callback_token_renewal_rate_limits.sql rename to apps/api/src/db/migrations/0180_workspace_callback_token_renewal_rate_limits.sql diff --git a/apps/api/src/services/session-sleep-snapshot-wait.ts b/apps/api/src/services/session-sleep-snapshot-wait.ts index c9f8d8a4ba..2b441dfd0c 100644 --- a/apps/api/src/services/session-sleep-snapshot-wait.ts +++ b/apps/api/src/services/session-sleep-snapshot-wait.ts @@ -4,7 +4,7 @@ import type * as schema from '../db/schema'; import type { Env } from '../env'; import { log } from '../lib/logger'; import { parsePositiveInt } from '../lib/route-helpers'; -import { hibernateAgentSessionOnNode,type HibernateCallbackTokenDelivery } from './node-agent'; +import { hibernateAgentSessionOnNode, type HibernateCallbackTokenDelivery } from './node-agent'; import { completeActiveSessionSnapshotAsDegraded, getSessionSnapshotCaptureState, diff --git a/tasks/archive/2026-10-04-workspace-callback-token-renewal.md b/tasks/archive/2026-10-04-workspace-callback-token-renewal.md index 0165648d42..ceec368791 100644 --- a/tasks/archive/2026-10-04-workspace-callback-token-renewal.md +++ b/tasks/archive/2026-10-04-workspace-callback-token-renewal.md @@ -112,7 +112,7 @@ Changes from local review (security, Go, Cloudflare, test, task-completion, docs `tasks/backlog/2026-10-04-instant-generation-aware-callback-token-renewal.md`. 7. **Rate limit (rule 28 §4).** Authenticated renewal attempts count against `RATE_LIMIT_CALLBACK_TOKEN_RENEWAL` (default 12 per `RATE_LIMIT_CALLBACK_TOKEN_RENEWAL_WINDOW_SECONDS` - = 3600) per workspace, in one guarded D1 upsert (migration 0179, cascade-deleted with the + = 3600) per workspace, in one guarded D1 upsert (migration 0180, cascade-deleted with the workspace). A slot is spent only after both proofs, the node binding and the active check, so a caller without the credentials cannot use up the legitimate agent's quota. 429 + `Retry-After`; the agent backs off. @@ -149,7 +149,7 @@ Changes from local review (security, Go, Cloudflare, test, task-completion, docs - [x] `node-agent-session-snapshots.ts`: include fresh token on hibernate when bound + active (VM only) - [x] Env vars documented: env-reference skill (API `CALLBACK_TOKEN_*`, agent `WORKSPACE_CALLBACK_TOKEN_*`, `MSG_AUTH_RENEWAL_WAIT`), public VM agent reference -- [x] Review: atomic per-workspace renewal limit (migration 0179, `workspace-callback-token-renewal-rate-limit.ts`, +- [x] Review: atomic per-workspace renewal limit (migration 0180, `workspace-callback-token-renewal-rate-limit.ts`, `RATE_LIMIT_CALLBACK_TOKEN_RENEWAL[_WINDOW_SECONDS]`, 429 + Retry-After) - [x] Review: renewal refuses Instant runtimes (`isInstantRuntimeBinding`, shared with delivery) - [x] Review: one delivered token per hibernate wait (`HibernateCallbackTokenDelivery`, a plain memo object