Skip to content

Paused-pane stale annotation is intermittent: two of three resurfacings arrived unannotated #1600

Description

@m0d7

Symptom

A scout deliberately idling with a declared paused: status keeps producing bare, unannotated stale: wakes, each costing a full handling turn.

Observed today on a long-lived advisory scout that is intentionally kept alive between captain follow-ups. Reconciled state is unambiguous:

$ bin/fm-crew-state.sh van-target-state-advisor
state: paused · source: status-log · idle by design, awaiting captain follow-ups ...

Last status line:

paused: idle by design, awaiting captain follow-ups on D4 (reopened) and the target-state diagram

Yet the drained wake records are plain:

1785775969	542	stale	default:wA:pQ	stale: default:wA:pQ
1785779616	549	stale	default:wA:pQ	stale: default:wA:pQ

Roughly one every few minutes, across the same declared pause.

Why this looks wrong

bin/fm-watch.sh has a dedicated absorber for exactly this case - handle_paused_stale (around line 327) - which is documented to absorb a stale pane under a declared paused:, throttle resurfacing to a long cadence, and when it does resurface, label the reason:

stale: $win (paused ${age}s, awaiting external - declared pause, rechecked on a long cadence not a wedge; confirm the wait still holds)

The wakes received carry none of that annotation, so they do not appear to be coming through the paused path at all. status_is_paused_or_captain_held (bin/fm-classify-lib.sh:135) looks correct against this status line, so the mismatch is somewhere between the predicate and the wake emission - I did not isolate it further.

Impact

Two costs, both real but neither destructive:

  1. Each bare stale: wake demands a handling turn under the supervision contract, so a deliberately-idle worker generates steady turn churn.
  2. More importantly, the annotation is what tells the handler "declared pause, not a wedge". Without it the wake is indistinguishable from a genuinely stuck worker, which is precisely the signal handle_paused_stale exists to preserve. The risk is a real wedge being dismissed as more of this noise, or a healthy paused worker being recovered unnecessarily.

Repro shape

  1. Spawn a scout instructed to stay alive and idle after reporting.
  2. Have it append a paused: line as its last status event.
  3. Leave it idle past the stale threshold and watch the drained wake records.

Expect: absorbed, or resurfaced on a long cadence with the paused annotation. Actual: repeated bare stale: wakes.

Note

Long-lived advisory scouts that stay resident between captain follow-ups are a legitimate and useful pattern, so the paused-absorption path matters more than it might look - it is the only thing keeping a resident worker from generating indefinite wake churn.

Activity

  1. m0d7 commented on Aug 3, 2026

    @m0d7
    Author

    Correction to this report - the original framing was wrong on the important part, and the defect is much narrower than filed.

    What I got wrong

    I claimed the bare wakes arrived "roughly one every few minutes" and that they "do not appear to be coming through the paused path at all". Both are wrong.

    The two records I cited are 3647 seconds apart - 61 minutes, not minutes apart. I misread the timestamps when filing. So the long-cadence throttle in handle_paused_stale is demonstrably working: the paused pane is being absorbed and resurfaced on a ~60-minute cadence exactly as designed, not spamming.

    A third resurfacing then arrived with the annotation present and correct:

    stale: default:wA:pQ (paused 3844s, awaiting external - declared pause, rechecked on a long cadence not a wedge; confirm the wait still holds)
    

    So there is no turn-churn problem and no missing absorption. The bulk of the original impact section does not apply.

    What appears to remain

    A narrower inconsistency: across three resurfacings of the same declared pause on the same pane, two carried no annotation and the third carried it in full. All three were on the same long cadence.

    1785775969  542  stale  default:wA:pQ  stale: default:wA:pQ
    1785779616  549  stale  default:wA:pQ  stale: default:wA:pQ
    1785783229  550  stale  default:wA:pQ  stale: default:wA:pQ (paused 3844s, awaiting external - declared pause, ...)
    

    The paused: status line was already the last status event before the first of these, so the pause was declared throughout.

    Why it is still worth a look, at lower priority

    The annotation is the only thing distinguishing "declared pause, not a wedge" from a genuinely stuck worker at the point of handling. When it is intermittently absent the handler loses that signal on those occurrences - so the risk is misclassification, not noise.

    Plausibly the bare ones precede the .paused-* marker being established for that pane hash, in which case the fix is ordering rather than logic. I did not isolate it, and I am no longer confident enough in my reading of the internals to assert a cause.

    Suggest re-triaging as a low-priority consistency fix, or closing if the first-resurfacing-after-pause case is intended.

    Apologies for the noise in the original report.

  2. changed the title [-]Declared paused: scout still emits repeated bare stale: wakes without the paused annotation[/-] [+]Paused-pane stale annotation is intermittent: two of three resurfacings arrived unannotated[/+] on Aug 3, 2026
  3. kunchenguid commented on Aug 27, 2026

    @kunchenguid
    Owner

    Speaking as Kun's firstmate: inspected on current main 4f89f5b5e235 (#3184). First look.

    Author's follow-up stands: the original "every few minutes / not the paused path" framing is wrong. The three cited records are ~3600s apart, so handle_paused_stale's long cadence is working. What remains is the missing annotation on two of those three resurfacings.

    VISION.md (each rule):

    • One captain, one interface: aligns. The annotation is the only handler-facing distinction between "declared pause, not a wedge" and a genuine stuck worker. Losing it on some long-cadence wakes is dishonest signal, even at low rate.
    • Authority is explicit: aligns. Not a new grant.
    • Scripts own the mechanics: aligns. Pause vs wedge is exact; the handler should not re-adjudicate a bare stale: that the watcher already classified as pause.
    • A restart is a non-event: aligns. A declared wait must keep reading as a declared wait across polls.
    • Delegation with a spine: aligns. Fix the existing paused-absorb path.
    • The fleet outlives any vendor: aligns. Backend-agnostic watcher classification.
    • Scope: aligns. Watcher mechanics.

    Evidence on 4f89f5b5e235: Still real. pause_state_class admits a live ordinary paused crew to paused only after an agent-liveness dead verdict (or for a secondmate). Otherwise it returns none. First sight of a hash then hits surface_nonterminal_stale, which emits bare stale: $win and only then writes .paused-*. Same-hash later polls go through handle_paused_stale (annotated, throttled). A hash change (or a watcher that re-sees the pane as new) repeats the bare first-sight. That matches two unannotated long-cadence wakes then one annotated one; it is not "first resurfacing is intended."

    Open #2738 (Omar-Nawaf, for #2614) is the help: pause_state_class admits a latest-line paused: to the long cadence on the declaration alone (no agent-liveness read), which routes those polls through handle_paused_stale instead of surface_nonterminal_stale. It does not cite #1600. Not opening a competing PR. Please cite Fixes #1600 (or equivalent) on #2738.

    Not labeling ready-for-pr.

  4. devin-ai-integration commented on Sep 24, 2026

    @devin-ai-integration
    No description provided.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions