What happens
bin/fm-crew-state.sh and the watcher's staleness triage read the LAST LINE of state/<id>.status to determine a worker's current state.
A worker that appends a legitimate status record as a MULTI-LINE block ends the file on ordinary prose. The last-line read then finds no state verb, so the worker reads as unknown / undeclared-idle, and the watcher escalates it as a possible wedge.
Why it matters
The failure is silent and it points the supervisor at the wrong thing. Observed on 2026-08-27:
- firstmate nudged one worker four times to "declare your wait"
- and told the captain twice that the worker had gone quiet
The worker had declared its wait every single time, as a multi-line paused: record. The fault was entirely in the read. The worker diagnosed it, not the supervisor.
Each false escalation costs a full model turn, and worse, it erodes the signal: a supervisor who is told four times that a healthy worker is wedged learns to discount the alarm.
Why fixing the reader beats fixing the workers
The brief scaffold can ask for single-line status records, and should. But that is a rule every worker must remember forever, enforced by nothing, and a multi-line append is a natural thing to write when the status carries detail. The reader is one place; the workers are unbounded.
Suggested shape, not prescriptive
Scan backwards from the end of the file for the most recent line that begins with a known state verb (working:, paused:, blocked:, needs-decision:, resolved:, done:, failed:, note:), rather than parsing only the final line. That also makes the read robust against any trailing blank line.
Consider whether bin/fm-classify-lib.sh's own triage shares the assumption; if it does, both should be fixed together so the watcher and the state reader cannot disagree.
Related
Same shape as an inverted log-severity defect found in a project on the same day: a reporting surface that misleads whoever reads it into diagnosing the wrong thing. Both cost real investigation time.
What happens
bin/fm-crew-state.shand the watcher's staleness triage read the LAST LINE ofstate/<id>.statusto determine a worker's current state.A worker that appends a legitimate status record as a MULTI-LINE block ends the file on ordinary prose. The last-line read then finds no state verb, so the worker reads as
unknown/ undeclared-idle, and the watcher escalates it as a possible wedge.Why it matters
The failure is silent and it points the supervisor at the wrong thing. Observed on 2026-08-27:
The worker had declared its wait every single time, as a multi-line
paused:record. The fault was entirely in the read. The worker diagnosed it, not the supervisor.Each false escalation costs a full model turn, and worse, it erodes the signal: a supervisor who is told four times that a healthy worker is wedged learns to discount the alarm.
Why fixing the reader beats fixing the workers
The brief scaffold can ask for single-line status records, and should. But that is a rule every worker must remember forever, enforced by nothing, and a multi-line append is a natural thing to write when the status carries detail. The reader is one place; the workers are unbounded.
Suggested shape, not prescriptive
Scan backwards from the end of the file for the most recent line that begins with a known state verb (
working:,paused:,blocked:,needs-decision:,resolved:,done:,failed:,note:), rather than parsing only the final line. That also makes the read robust against any trailing blank line.Consider whether
bin/fm-classify-lib.sh's own triage shares the assumption; if it does, both should be fixed together so the watcher and the state reader cannot disagree.Related
Same shape as an inverted log-severity defect found in a project on the same day: a reporting surface that misleads whoever reads it into diagnosing the wrong thing. Both cost real investigation time.