Skip to content

fix: log a listen task that never connects - #46

Merged
Bre77 merged 4 commits into
mainfrom
fm/tstream-silence-watchdog
Sep 24, 2026
Merged

Bre77 merged 4 commits into
mainfrom
fm/tstream-silence-watchdog

Conversation

@Bre77

@Bre77 Bre77 commented Sep 14, 2026 •

Copy link
Copy Markdown
Member

Intent

A customer's Home Assistant instance restarted and its listen task never logged a single connect attempt for 2.5+ days - zero server-side connect requests, while every ordinary REST call from the same instance went through fine. The exact cause isn't provable after the fact, but the failure mode itself was silent: listen() could return or raise before ever reaching a successful connect with nothing in the caller's logs to show it. listen() now tracks whether it has seen a successful connect and, if it exits before that, logs one WARNING line describing the failure before re-raising - nothing before a first connect is swallowed.

…r connects

Add a silence watchdog in __anext__: each SSE read is bounded by a single
30s window (three missed 10s keepalives), so a connection that stops
producing any bytes - not just a hard error - gets closed and picked up by
the existing reconnect loop. connect()'s sock_read is left unbounded so the
two timeouts don't race. listen() now tracks whether it has ever reached a
successful connect; if it returns or raises before that, it logs at ERROR
and re-raises rather than dying quietly.

Motivated by a customer's Home Assistant instance whose listen task never
logged a single connect attempt for 2.5+ days after a restart, with no
server-side or client-side trace of why.
@Bre77 Bre77 added the fm Opened by a Firstmate crewmate label Sep 14, 2026
@Bre77
Bre77 marked this pull request as draft September 14, 2026 22:42
…overhead

Wrapping each SSE line read in asyncio.wait_for (the silence watchdog) adds
a couple of extra event-loop hops before a completed read is observable on
Python <3.12, where wait_for still allocates a bridging waiter future
instead of the 3.12+ fast path. Several existing tests advanced the loop by
a fixed, version-tuned number of sleep(0) ticks that no longer covered this;
give them enough margin to pass on every supported Python version. No
production behavior changes - the clean-end path still logs INFO and
reconnects immediately.
…e watchdog

sock_read=30 on the connect() request already bounds a connection that
stops producing bytes (data or SSE keepalive) and routes into the
existing ClientError handler, which already logs and reconnects - the
added wait_for-based watchdog was redundant complexity for no new
behavior. Restore sock_read=30 and drop the wait_for wrapper and streak
logging. Keep only the early-exit change: listen() logs one WARNING
line if it returns or raises before its first successful connect, and
still re-raises rather than swallowing the failure.
@Bre77 Bre77 changed the title fix: reconnect a silently stalled stream, log a listen task that never connects fix: log a listen task that never connects Sep 14, 2026
Keep the PR diff scoped to the listen() early-exit change only.
@Bre77
Bre77 marked this pull request as ready for review September 24, 2026 07:24
@Bre77
Bre77 merged commit 0e891c4 into main Sep 24, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

fm Opened by a Firstmate crewmate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant