Skip to content

A declared 'paused:' external wait still wedge-escalates on aging, reaching demand-deep-inspection on a provably healthy lane #3909

Description

@ruant

Observed

A crewmate declared a bounded external wait and then re-escalated as a possible wedge three times in ~15 minutes, reaching demand-deep-inspection, while it was demonstrably healthy the whole time.

Task fix-clock-dependent-tests, 2026-09-07. Its last status line was:

paused: final validation at step 6/6 - clean whole-assembly baseline (~20 min); steps 1-5 all green

The wakes received, on the same pane:

stale: default:w5:pMN (idle 251s, possible wedge, escalation 1)
stale: default:w5:pMN (idle 288s, possible wedge, escalation 2)
stale: default:w5:pMN (idle 300s, possible wedge, escalation 3, demand-deep-inspection: same pane has wedge-escalated 3 times in a row - do not re-absorb on the run-step/pane state alone)

At the third escalation the lane was verifiably working:

  • bin/fm-crew-state.sh fix-clock-dependent-tests -> state: working · source: pane · harness busy (claude-hook)
  • a testhost.dll process for that worktree at 169% CPU, 2m40s elapsed
  • its final-assembly.log written 30 seconds earlier, mid test run

So all three escalations were on a healthy lane doing exactly what it had declared.

Why this costs something

demand-deep-inspection explicitly instructs the handler not to re-absorb on run-step or pane state alone, which is the correct instruction for a genuine repeat-wedge candidate. Here it forced a full process-and-log inspection of a lane whose own declared wait had not yet elapsed. Each escalation consumes a firstmate turn.

The cost scales with lane count. With nine lanes sharing a four-slot build semaphore, long queue waits and 20-minute suite runs are the normal case, not the exception, so healthy lanes generate a steady stream of wedge escalations.

Relevant code, without claiming the fix

bin/fm-classify-lib.sh crew_absorb_class returns paused for a declared external wait and working for a busy pane, and its own comment says callers run it "ONLY on no-verb signal and first-sighting stale paths, never every wake". Aging past FM_STALE_ESCALATE_SECS (default 240s) therefore appears to escalate without re-consulting that verdict. That reading is from the source and the observed wake sequence; I have not traced the caller, so treat the mechanism as a starting point rather than a diagnosis.

There is also a documented promise worth checking against behaviour. The generated crewmate brief tells the worker that paused: means "firstmate then leaves your idle pane alone and rechecks it on a long cadence instead of treating it as a possible wedge". A declared pause that still escalates three times in 15 minutes does not match what the worker was told.

What would resolve it

Either the aging path consults the declared-pause verdict before escalating, or the brief's promise is corrected to describe what actually happens. Both are defensible; the current state is that the two disagree, which is the harder thing to reason about.

Filed from a live fleet run rather than a code read, so the evidence above is the observation, not a reproduction script.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ready-for-prTriage: real bug or VISION-aligned feature, open for a PR

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions