Observed on main at f66be0f.
Symptom
During an away-mode stretch, the sub-supervisor produced 258 and then 66 consecutive identical check: rearm-resurface escalations, back to back.
state/.supervise-daemon.watcher.err showed 61 consecutive watcher cycles with identical output and nothing else.
Each cycle is a captain-facing escalation, so the practical effect is an unusable away-mode session.
Cause
In bin/fm-supervise-daemon.sh, the watcher restart loop classifies a child exit three ways.
rc != 0 or an empty reason calls record_crash and sleeps backoff_secs.
- A non-wake stdout line sleeps
HOUSEKEEPING_TICK.
- A valid wake reason logs it, handles durable wakes, and calls
start_watcher immediately, with no sleep at all.
The third path has no throttle of any kind.
That is correct for normal fleet traffic, where consecutive wakes differ.
But when the watcher keeps exiting with the same wake reason - for example an unconsumed recovery-marker episode kept alive by ongoing wake-queue traffic - the loop spins as fast as the watcher can start and exit.
The adjacent non-wake branch already carries a comment describing exactly this failure mode:
Non-wake stdout ... is NOT a wake: idling here prevents an escalation flood and a backoff-less child restart.
So the hazard is already recognized in the code, just not guarded for a repeating valid wake.
Suggested shape, not a prescription
Tracking consecutive occurrences of an identical wake reason and reusing the existing threshold, window, and backoff values would throttle it without calling it a crash, since it genuinely is not one.
We are running a local patch along those lines and it resolves the flood, but we are deliberately not opening a PR.
You will have a much better sense of whether this is the right layer for the fix.
Happy to share the patch if it is useful.
Related
The broken-pipe line reported separately was present alongside every one of these cycles and looked like the cause for a long time.
It is not.
Observed on
mainatf66be0f.Symptom
During an away-mode stretch, the sub-supervisor produced 258 and then 66 consecutive identical
check: rearm-resurfaceescalations, back to back.state/.supervise-daemon.watcher.errshowed 61 consecutive watcher cycles with identical output and nothing else.Each cycle is a captain-facing escalation, so the practical effect is an unusable away-mode session.
Cause
In
bin/fm-supervise-daemon.sh, the watcher restart loop classifies a child exit three ways.rc != 0or an empty reason callsrecord_crashand sleepsbackoff_secs.HOUSEKEEPING_TICK.start_watcherimmediately, with no sleep at all.The third path has no throttle of any kind.
That is correct for normal fleet traffic, where consecutive wakes differ.
But when the watcher keeps exiting with the same wake reason - for example an unconsumed recovery-marker episode kept alive by ongoing wake-queue traffic - the loop spins as fast as the watcher can start and exit.
The adjacent non-wake branch already carries a comment describing exactly this failure mode:
So the hazard is already recognized in the code, just not guarded for a repeating valid wake.
Suggested shape, not a prescription
Tracking consecutive occurrences of an identical wake reason and reusing the existing threshold, window, and backoff values would throttle it without calling it a crash, since it genuinely is not one.
We are running a local patch along those lines and it resolves the flood, but we are deliberately not opening a PR.
You will have a much better sense of whether this is the right layer for the fix.
Happy to share the patch if it is useful.
Related
The broken-pipe line reported separately was present alongside every one of these cycles and looked like the cause for a long time.
It is not.