Skip to content

Sub-supervisor restarts the watcher with no backoff when it keeps returning the same wake reason #3274

Description

@vishalvx

Observed on main at f66be0f.

Symptom

During an away-mode stretch, the sub-supervisor produced 258 and then 66 consecutive identical check: rearm-resurface escalations, back to back.
state/.supervise-daemon.watcher.err showed 61 consecutive watcher cycles with identical output and nothing else.
Each cycle is a captain-facing escalation, so the practical effect is an unusable away-mode session.

Cause

In bin/fm-supervise-daemon.sh, the watcher restart loop classifies a child exit three ways.

  • rc != 0 or an empty reason calls record_crash and sleeps backoff_secs.
  • A non-wake stdout line sleeps HOUSEKEEPING_TICK.
  • A valid wake reason logs it, handles durable wakes, and calls start_watcher immediately, with no sleep at all.

The third path has no throttle of any kind.
That is correct for normal fleet traffic, where consecutive wakes differ.
But when the watcher keeps exiting with the same wake reason - for example an unconsumed recovery-marker episode kept alive by ongoing wake-queue traffic - the loop spins as fast as the watcher can start and exit.

The adjacent non-wake branch already carries a comment describing exactly this failure mode:

Non-wake stdout ... is NOT a wake: idling here prevents an escalation flood and a backoff-less child restart.

So the hazard is already recognized in the code, just not guarded for a repeating valid wake.

Suggested shape, not a prescription

Tracking consecutive occurrences of an identical wake reason and reusing the existing threshold, window, and backoff values would throttle it without calling it a crash, since it genuinely is not one.
We are running a local patch along those lines and it resolves the flood, but we are deliberately not opening a PR.
You will have a much better sense of whether this is the right layer for the fix.
Happy to share the patch if it is useful.

Related

The broken-pipe line reported separately was present alongside every one of these cycles and looked like the cause for a long time.
It is not.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingready-for-prTriage: real bug or VISION-aligned feature, open for a PR

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions