fix(dsh-plugin): reconnect the observation stream after a fatal failure - #315
Conversation
The observation store attached only an onmessage handler to the plugin's SSE stream and start() returns early while the store is already marked started, so a fatal EventSource failure left the view empty until the page was reloaded. An EventSource never retries a non-200 response, which is exactly what happens when the page loads before the observation route is registered, or right after a plugin reload: the console reports 404 on /bsk-observation/events while /bsk-observation/state and a direct SSE request both answer normally, and the sidebar sits at "no session" with no way to recover except a page reload. - attach an onerror handler and rebuild the stream with a bounded backoff when readyState is CLOSED; transient drops keep readyState 0/1 and are still left to the EventSource's own retry - reset the backoff on any healthy frame, and cancel a pending reconnect in stop() - republish the snapshot whenever the stream is (re)created so `subscribed` reflects reality - surface `subscribed` in the overlay and sidebar empty state: a dead feed now reads "connecting..." / "connection lost - retrying" instead of "no session" - keep the floating card mounted while reconnecting (it used to return null as soon as there were no sessions, showing nothing at all)
|
Thanks — the failure mode is real, and reconnecting only when Two things need to change before this is ready.
Please also cover both cases in tests: a recreated stream that does not immediately emit a snapshot, and |
`subscribed` is derived from whether an EventSource object exists, so it stays true for the whole outage: the failed stream is still held until the rebuilt one replaces it. The "connecting…" states therefore never showed during a real failure, while the floating card, now gated on `subscribed`, was committed for one frame with "connection lost — retrying" on every page load that had no sessions. - add `reconnecting` to the snapshot: set on any stream error, including transient drops the browser retries itself, and cleared by the next well-formed frame on the current stream - show "reconnecting…" in the overlay and sidebar header while it is set, in place of the last session status that could no longer be trusted - keep the floating card hidden without sessions, as before this change - cancel a pending reconnect whenever the stream is rebuilt, so a thumbnail switch during the backoff is not torn down by the stale timer - drop the extra publish per stream rebuild that only served `subscribed`
The expanded header already replaced a stale "s1 · clicking" line with "reconnecting…", but the collapsed capsule and both status dots still read the last session as live. Collapse is the form users leave on the page, so an outage there looked like an in-progress action. - treat reconnecting as its own chrome state so the header and capsule dots leave the active color - keep the capsule's session count and swap the action timer for "reconnecting…" until the feed delivers a frame again
The observation store attached only an onmessage handler to the plugin's SSE stream and start() returns early while the store is already marked started, so a fatal EventSource failure left the view empty until the page was reloaded.
An EventSource never retries a non-200 response, which is exactly what happens when the page loads before the observation route is registered, or right after a plugin reload: the console reports 404 on /bsk-observation/events while /bsk-observation/state and a direct SSE request both answer normally, and the sidebar sits at "no session" with no way to recover except a page reload.
subscribedreflects realitysubscribedin the overlay and sidebar empty state: a dead feed now reads "connecting..." / "connection lost - retrying" instead of "no session"Follow-up (d102adc)
The reconnect logic above is kept as is. Review found that
subscribedstays true for the whole outage, because the failed EventSource is still held until the rebuilt one replaces it. The "connecting…" states therefore never showed during a real failure, while the floating card, gated onsubscribed, was committed for one frame on every page load with no sessions. This commit:reconnectingto the snapshot: set on any stream error (including transient drops the browser retries itself) and cleared by the next well-formed frame on the current stream;subscribedkeeps its original meaningmain; the sidebar tab still reports the outagesubscribedTests cover the outage state across fatal and transient drops, the thumbnail switch during a pending reconnect, the header text while reconnecting, and that the card is never committed without sessions. Checked on the branch and merged onto current
main: typecheck,pnpm test(376 passed after the merge), biome, and stylelint are all clean.Follow-up (8f485ab)
The expanded header now says "reconnecting…", but the collapsed capsule and both status dots still treated the last session as live. Collapse is the form users leave on the page.
reconnectingis its own chrome state, so the header and capsule dots leave the active colorA hard ceiling on retries (permanent 403 / gone route) is left as a later change. Delay already caps at 30s; each store then costs about two local requests a minute. A "give up and retry now" surface would add new store and UI state without changing the recovery this PR is for.